Fit Finder: Redesigning trust, not just size
I led the product design evolution of Fit Finder across two companies. First at Fit Analytics, then after Snap's acquisition in 2021. The redesign shipped in 2023 to clients including Zara, Hugo Boss, and ASOS, serving millions of shoppers monthly.
Completion rate*
Conversion rate*
Return rate*
*Three-month A/B test, V5.0 against V4.2, across 20 retail partners.
My role
Design lead, Fit Finder V5.0
Timeline
Five months (2022–2023)
Team
PM, User Researcher, Data Scientists, Engineers
Platform
Web & Mobile (white-label SDK)
Context
A product I knew from the inside
Fit Finder is a white-label size recommendation tool embedded directly into retailer product pages. When a shopper clicks "Find my size," Fit Finder walks them through a short questionnaire about their body measurements and fit preferences, then surfaces a personalised size recommendation. The goal: reduce the uncertainty that drives returns, and give shoppers the confidence to buy.

Fit Finder (V4.2), Upper body / Woman — Mobile
I joined Fit Analytics in 2019 as a senior designer, working on the product long before Snap acquired the company in 2021. After the acquisition, Fit Finder became part of Snap's ARES Shopping Suite, and I continued leading design as we planned a major version update. That continuity mattered. I wasn't parachuted in to redesign something I'd never used. I knew what the data said, where users dropped off, and which parts of the product engineering had quietly flagged as technical debt.

Fit Finder (V4.2), Upper body / Woman — Desktop
The V5.0 redesign was the most significant project in a four-year tenure on the product. Footwear ran as a parallel track: we cut several screens and rebuilt the UI on others, and it went into another A/B test. This case study follows the upper and lower body work, where the research went deepest and the before-and-after is clearest.
Problem
What was actually breaking
V4.2 had a measurable drop-off problem. Analytics showed significant abandonment at the brand-selection screen; a step that asked users to select their preferred brand after entering body measurements. The intent was to calibrate recommendations to brand-specific sizing, but the data science team had been questioning its value for months.
Beyond that, usability testing pointed to a broader trust gap. Users were completing the flow but not acting on the recommendation. They'd see their size and then still check the size chart manually, or abandon the purchase. The interface felt bureaucratic; too many screens and inconsistent visual language across clients.
The core problem statement
Only 14% of shoppers who opened Fit Finder ever reached a size recommendation, and the brand screens were the steepest drop after the entry screen. How might we get more people to a recommendation, and make sure the one they reach is worth the trip, without losing accuracy?
Research & Analysis
What the research actually told us
Two weeks of research before touching a single design file. Working alongside a dedicated User Researcher and the Data Science team, we ran interviews, usability testing, empathy mapping, journey mapping, analytics review, card sorting, and a similarity matrix. The goal was to build enough evidence to make decisions we could defend.
Who we were designing for
We started by defining three user archetypes from interview synthesis, then used them throughout the project to pressure-test design decisions. They weren't fictional personas. They were distilled from real patterns we heard across sessions.

The Confident Buyer
Knows her size, wants confirmation
Mia knows she's a medium. She uses Fit Finder to validate, not discover. If the recommendation contradicts her expectation without explanation, she ignores it and checks the size chart anyway. The trust gap on the recommendation reveal screen was her problem.

The Uncertain Measurer
Wants help, gets stuck on inputs
Omar knows his height and weight, but the unit toggle between metric and imperial gives him pause mid-flow. It's a small friction point, but for a user already uncertain about trusting the recommendation, doubt at the input stage compounds into doubt at the result.

Three archetypes built from interview synthesis — each one maps to a specific failure point in the V4.2 flow
Where the journey broke down
Journey mapping the V4.2 flow made the problem visible in a way that analytics alone couldn't. We mapped the emotional arc across every screen; from the moment a shopper clicked "Find my size" through to the recommendation reveal, and marked where confidence dropped, where confusion spiked, and where users abandoned the flow entirely.
Two moments stood out. The brand selection screen caused a sharp confidence dip for users whose brand wasn't listed, and the recommendation reveal triggered disbelief for users whose result didn't match their expectation. Neither moment had any recovery mechanism in V4.2.

Usability test findings mapped to each screen in the V4.2 flow
What the usability tests found, screen by screen
We ran moderated usability sessions with 12 participants across the full V4.2 flow, conducted remotely via Google Meet. The findings were specific enough to drive direct design decisions, not just general friction, but screen-by-screen evidence of where and why users lost confidence. The clearest was the progress indicator: 8 of 12 testers could not tell how far through the flow they were, which is why the ellipses came out in V5.0.
What empathy mapping revealed about trust
Empathy mapping sessions after the usability tests and interviews helped us understand what users were thinking and feeling at the moments that mattered most. The pattern was consistent: when a recommendation contradicted a user's expectation, they immediately looked for a reason and found nothing.
The product did explain itself, just not at the moment it mattered. There was a confidence figure and a line about what similar shoppers bought, but it sat under the recommendation in small text, it was the same sentence for everyone, and it never addressed the one thing the shopper had brought with her: the size she already believed she was. A generic explanation answers a question nobody asked. What they wanted to know was not "how does this work", it was "why is this different from what I expected".
I know I'm usually a medium, but this thing told me large. I didn't really trust it, so I just guessed anyway.
— Usability test participant, 2022

V4.2 funnel across the 20 retail partners later used in the A/B test. 100% is shoppers who opened Fit Finder. A third leave on the first measurement screen before entering anything. Of those who stay, the steepest single drop is Brand selection, and completion keeps falling across every subsequent brand screen.
What analytics and the similarity matrix confirmed
Analytics review revealed a clear pattern. Shoppers who made it through measurements, body shape and personal data abandoned at a disproportionate rate once they hit the brand screens. The steepest single drop after the entry screen was Fit preference at 49% to Brand selection at 36%, a 13 point fall in one step, and completion kept eroding across three more brand screens down to 20%.
It did not stop there. Only 14% ever reached a result: another six points lost at the final step, among shoppers who had already answered every question. Whether that was fatigue or the load time before the reveal, we never isolated it. What the qualitative work had already told us was that reaching a result and acting on one were different things. Two problems in one funnel, and they set the agenda between them: fewer screens to reach a recommendation, and a recommendation worth reaching.

Card sorting results: shoppers grouped height, weight, body shapes and bra size together as body data, kept age and gender as a separate pair, and placed brand preference in a category of its own.
Card sorting provided the mental model evidence to support removing it. Shoppers grouped height, weight, body shapes and bra size together as body data, kept age and gender as a separate pair, and placed brand selection in a category of its own. We were running three mental models through one linear flow.
The similarity matrix made that pattern quantitative. With 10 respondents, the blue clusters show near-perfect agreement; height, weight, chest shape, belly shape, hip shape, and bra size were grouped together by almost everyone. Brand information scored zero across those same groupings. The data wasn't ambiguous: brand data belonged to a mental model of its own, and the V4.2 flow had collapsed three into one.

Similarity matrix from the card sorting exercise — blue clusters show how users grouped personal measurements together and separated brand information into a distinct category
01
Brand data: no accuracy impact
The Data Science team's model analysis found brand selection made no measurable difference to accuracy. Four screens of user effort, no gain in the output.
02
Three mental models, one linear flow
Card sorting separated body data (measurements, shapes, bra size) from demographics (age, gender) from brand preference. V4.2 ran all three through one sequence as if they were the same question.
03
5 of 12 testers could not find their brand
Usability data showed 5 of 12 testers couldn't locate their brand in the selection screen. Either it wasn't listed, or they couldn't find it in "More Brands." Drop-off followed immediately.
04
4 of 12 testers ignored the recommendation
Post-flow interviews confirmed 4 of 12 testers didn't act on the size recommendation. They second-guessed it and checked the size chart manually. The problem wasn't the model, it was the interface.
Together these pointed one way. The brand screens were costing us a third of the funnel and returning nothing to the model.
Design
From sketches to a system
Six weeks of design work, moving from Crazy 8s through sketches, wireframes, and high-fidelity screens. The process ran alongside an accessibility audit and regular design critiques. I also involved junior designers throughout as a way to both pressure-test decisions and develop the team's skills in parallel.
Evaluated the existing V4.2 for heading hierarchy, focus order, screen reader compatibility, and WCAG colour contrast. Identified accessibility issues across the existing flow that needed resolution before any visual redesign.
Fast ideation on alternative approaches to the measurement input, recommendation reveal, and trust-building moments. Eight concepts in eight minutes, then down-selection based on feasibility and user insight alignment.
Mapped the simplified flow without brand screens. Validated with the PM and engineering lead before moving to high fidelity. No wasted polish on concepts that wouldn't ship.
Designed within, and contributed to, the ARES Shopping Suite design system. Built reusable components for measurement inputs, recommendation cards, and progress indicators that could scale across Snap's AR tools.
Interactive prototype tested before handoff. The reveal screen dropped the competing size bars for a single recommendation, and added fit context for garments that run large or small.
From research to decisions — annotated V5.0 screens
The V5.0 recommendation reveal screen addressed the trust gap directly. One size rather than two competing options, a plain-language line grounded in purchase behaviour at scale, and a fit note for garments that run large or small.

Every change in V5.0 was driven by a specific finding from usability testing; brand screens removed, radio buttons replaced with buttons, arrows and ellipses stripped from body shape screens, and help text added to the bra size screen to reduce discomfort
Design system contribution
Fit Finder V5.0 was designed within the ARES Shopping Suite design system; a shared component library built to unify Snap's AR tools including virtual try-on and interactive product displays. I created the sizing-specific components (measurement inputs, recommendation cards, fit preference selectors) so they could be reused across future ARES products without redesign.

Fit Finder components built into the ARES Design System for reuse across Snap's Shopping Suite
Results
Three months, one decision
V5.0 launched in Q1 2023. Over a three-month A/B test across 20 retail partners, including global clients in apparel, the PM and Data Science team tracked completion, conversion, returns and satisfaction, with three of those set as rollout criteria.
Before the test began, the team aligned on three conditions for a full rollout decision: statistical significance at 95% confidence on conversion rate; no regression on return rate; and a neutral or positive shift in user satisfaction. All three were met, with return rate clearing its bar comfortably.
Completion rate
Conversion rate
Return rate
Each number alone could mislead. Completion alone means more people saw an answer, not that they believed it. A conversion uplift can mean people were persuaded into things that don't fit. A return rate drop can mean fewer people bought at all. Measured together across the same 20 partners, they describe one chain: more shoppers reached a recommendation, more of them acted on it, and fewer sent the item back. Completion was the figure the problem statement opened on.
Removing the brand screens was the largest single reduction in drop-off, because it deleted the steps where the losses were happening. The rebuilt reveal, with a plain explanation and a fit note for garments that run large or small, was flagged in qualitative feedback as a meaningful trust signal. Two decisions, one from research and one from prototype testing, pulling the same way.

V5.0 variant, women's upper body, across the 20 A/B partners. Six screens instead of ten.
With all three success criteria met and statistical significance confirmed, the PM and engineering teams moved to full rollout across all retail partners. V5.0 stayed the baseline for the rest of my time on the product, and the front door problem the funnel had been flagging since 2022 became the brief for the next cycle.
Reflection
With more time
A third of shoppers left on the very first screen, before entering a single measurement. It is the single largest loss anywhere in the funnel and we scoped it out, because it is an entry point and value proposition problem rather than a flow one. That was the right call for this release. It was the wrong thing to still be true two versions later.
Cutting the brand screens was right on the numbers, but one group had asked for brand-specific sizing and got nothing back. We argued garment-level fit context covered the same need without the work, and the aggregate supported it. We never segmented the test by shoppers who used to pick a brand. An aggregate win can hide a segment going backwards.
I audited V4.2 up front, which told us what was already broken. I never re-ran those checks on the new screens while they were still wireframes. Contrast and touch target issues surfaced at final review instead, so decisions had to be revisited late and small inconsistencies got into the system.
The test told us V5.0 beat V4.2 on every metric we tracked. It could not tell us how much came from the brand screens, how much from the reveal, and how much from the general cleanup, because we shipped them as one variant. I'd want an isolating cell: brand screens removed, nothing else.
Contributing to ARES while simultaneously shipping Fit Finder meant some components were built for the project first and generalised after. A few of them weren't quite right for other contexts and needed rework. A clearer split between project-specific and system-level work from the start would have saved that.
Role and collaborators
My role and the wider team
I was design lead on the project. That meant the accessibility audit, wireframes and high-fidelity design, the design system contribution, prototyping and developer handoff. On research I wrote the plan with our User Researcher and ran the synthesis with her; she facilitated the interviews and moderated the usability sessions. The decision to remove the brand screens was mine to push for, and I validated it with the PM and Data Science team before it touched a design file.
Engineering Lead + Developers
Owned implementation, managed the developer handoff process, and ran QA across web and mobile.
Data Scientists
Ran analytics review, validated the recommendation model changes, and led A/B test analysis.
User Researcher
Facilitated the interview programme, ran empathy mapping sessions, and moderated usability testing.
Product Manager
Scoped the project, aligned stakeholders across Snap and the retail clients, and oversaw the A/B test.













