Building a Net-New Provider Search Tool
I was impacted by a layoff before this feature shipped. The outcomes and metrics detailed in this case study represent the projected targets defined during the design phase.
Context
Provider directory data degrades rapidly across the healthcare industry. Practice moves, provider retirements, and shifting network affiliations happen constantly, while backend directories often update on a delayed monthly cycle. For users, a provider search that was accurate weeks ago can easily lead to incorrect information today.
The organization I worked for routed active members to a third-party platform to search for in-network care. Over time, member complaints mounted as users received unexpected out-of-network bills, despite having verified coverage on the platform prior to their appointment.
"I found my doctor on your website to confirm that she was in the network. Now I have a bill for $340 because apparently she's not. I feel duped."
— Member, enrolled 14 months, first support contact
After years of relying on underperforming third-party tools, the business secured the resources to build a proprietary solution. As lead UX designer, I partnered with a senior architect and three engineers to complete early-phase designs in eight weeks.
As a 0-to-1 build without a legacy product to reference, the challenge required defining the core problem space just as much as designing the end-to-end solution.
Problem
Stale data was only one barrier to trust. The original search experience hindered users from filtering by critical criteria such as care quality, specialization, and cost. Unresponsive filters and incomplete results directly threatened timely access to care.
Quantitative signals backed this up: search-related support tickets were climbing, and out-of-network penalty disputes were on a steady rise.
My goal was to identify the root causes behind these trends and surface any contributing systemic issues beyond data freshness.
Process
Research synthesis. Starting with raw inputs from support ticket data, I inserted summary text, internal notes, and coordinating category tags into an LLM for a first-pass clustering of complaint types (data accuracy vs. UI/findability vs. terminology confusion). The LLM's reasoning mode was chosen over speed in this instance. And to quality check, I then validated those clusters manually against a sample of real tickets to make sure the AI's categories actually held up against real language and context. Estimated time savings was at least a day to a day and a half.
From my findings it appeared members were running into three recurring failure modes:
- Provider directory staleness: As stated earlier, stale or inaccurate legacy data lead to surprise out-of-network bills
- Filter inconsistency: No way to filter or search by specific plan tier, so results didn't reflect what a member's actual plan covered
- Confusing terminology: "in-network," "preferred," and "tiered" providers were used with little explanation, and members didn't understand the difference or why it mattered financially
Key insights showed how Zocdoc diverged from legacy databases by relying on providers to maintain their insurance credentials for real-time verification, enabling high-trust filters like "accepts your insurance."
Meanwhile, platforms like Healthgrades built confidence with real-time trust signals ("insurance verified today"), elevated social proof, and eliminated dead-end searches through auto-adjusting radius settings.
Navigating constraints: Open and ongoing communication with the Senior Architect and the new back-end vendor provided insight into addressing data staleness moving forward. I knew that design enhancements weren't going to mask the flaw entirely. It was a matter of mitigating where possible, then designing honestly without offering guarantees of accuracy.
Trade-offs were a constant reality. Given the financial penalties for missing the launch date, feature enhancements like ratings and reviews were scheduled post-MVP.
Rapid prototyping. Using Figma Make, I generated UI variants to explore confidence indicators, quick-filter approaches, and copy directions for narrow-network provider groups ("preferred" vs. "direct"). At least a half a day or more was saved by collaborative AI usage.
I then leveraged UserTesting for rapid feedback across user segments, ensuring each iteration advanced in addressing real user needs. Qualitative and quantitative feedback for task success rates, satisfactory ratings, and general sentiment towards aesthetic and content organization informed design choices.
Automating these early explorations shifted my focus from manual asset creation to evaluating strategic trade-offs.
Key decisions & trade-offs
Design system vs. custom components
I had spent ample time building out a design system and component library, and the effort was yielding results in shortened design to development cycles. Given the nature of this build and its specialized interactions, I created multiple custom components to fit these unique use cases and customer needs.
Treat search as exploration, not data entry
Obviously the search functionality was a make-or-break feature to the project. As development progressed I noted it was being engineering for more exacting data entry. I pushed back, advocating for open-ended discovery to aid users in finding what they want when they do not know exact terms.
What I deprioritized.
Metadata tags like "short wait times" were positively received as trust signals in user testing. However, those data sources showed various levels of accuracy. Tags more prone to inaccuracy or like "highly rated bedside manner" or "online intake" were dropped in order to preserve user confidence.
Expected impact
Since I left before launch, these were the metrics the team and I agreed would indicate success, along with the validation plan defined pre-launch:
Reflection
Working in a compliance-sensitive domain clarified where AI tools were genuinely useful — compressing research synthesis and generating divergent prototype directions quickly — and where they weren't a substitute for domain judgment, particularly around anything touching legal risk or member trust.
If I had stayed through launch, the two things I most wanted to watch were [fill in — e.g. "whether the confidence indicator actually reduced surprise-billing complaints, and whether the plan-tier selection step caused any drop-off in search completion"], because [why those specifically mattered].