Case study, Fit Finder, Fit Analytics → Snap, 2019-2023

Case study, Fit Finder, Fit Analytics → Snap Inc, 2019-2023

Fit Finder: Redesigning trust, not just size

I led the product design evolution of Fit Finder across two companies. First at Fit Analytics, then after Snap's acquisition in 2021. The redesign shipped in 2023 to clients including Zara, Hugo Boss, and ASOS, serving millions of shoppers monthly.

Completion rate*

14% → 39%

14% → 39%

Conversion rate*

+14%

+14%

Return rate*

–12%

–12%

*Three-month A/B test, V5.0 against V4.2, across 20 retail partners.

My role

Design lead, Fit Finder V5.0

Timeline

Five months (2022–2023)

Team

PM, User Researcher, Data Scientists, Engineers

Platform

Web & Mobile (white-label SDK)

Context

A product I knew from the inside

Fit Finder is a white-label size recommendation tool embedded directly into retailer product pages. When a shopper clicks "Find my size," Fit Finder walks them through a short questionnaire about their body measurements and fit preferences, then surfaces a personalised size recommendation. The goal: reduce the uncertainty that drives returns, and give shoppers the confidence to buy.

Fit Finder (V4.2), Upper body / Woman — Mobile

I joined Fit Analytics in 2019 as a senior designer, working on the product long before Snap acquired the company in 2021. After the acquisition, Fit Finder became part of Snap's ARES Shopping Suite, and I continued leading design as we planned a major version update. That continuity mattered. I wasn't parachuted in to redesign something I'd never used. I knew what the data said, where users dropped off, and which parts of the product engineering had quietly flagged as technical debt.

Fit Finder (V4.2), Upper body / Woman — Desktop

The V5.0 redesign was the most significant project in a four-year tenure on the product. Footwear ran as a parallel track: we cut several screens and rebuilt the UI on others, and it went into another A/B test. This case study follows the upper and lower body work, where the research went deepest and the before-and-after is clearest.

Problem

What was actually breaking

V4.2 had a measurable drop-off problem. Analytics showed significant abandonment at the brand-selection screen; a step that asked users to select their preferred brand after entering body measurements. The intent was to calibrate recommendations to brand-specific sizing, but the data science team had been questioning its value for months.

Beyond that, usability testing pointed to a broader trust gap. Users were completing the flow but not acting on the recommendation. They'd see their size and then still check the size chart manually, or abandon the purchase. The interface felt bureaucratic; too many screens and inconsistent visual language across clients.

The core problem statement

Only 14% of shoppers who opened Fit Finder ever reached a size recommendation, and the brand screens were the steepest drop after the entry screen. How might we get more people to a recommendation, and make sure the one they reach is worth the trip, without losing accuracy?

Research & Analysis

What the research actually told us

Two weeks of research before touching a single design file. Working alongside a dedicated User Researcher and the Data Science team, we ran interviews, usability testing, empathy mapping, journey mapping, analytics review, card sorting, and a similarity matrix. The goal was to build enough evidence to make decisions we could defend.

Who we were designing for

We started by defining three user archetypes from interview synthesis, then used them throughout the project to pressure-test design decisions. They weren't fictional personas. They were distilled from real patterns we heard across sessions.

The Confident Buyer

Knows her size, wants confirmation

Mia knows she's a medium. She uses Fit Finder to validate, not discover. If the recommendation contradicts her expectation without explanation, she ignores it and checks the size chart anyway. The trust gap on the recommendation reveal screen was her problem.

The Uncertain Measurer

Wants help, gets stuck on inputs

Omar knows his height and weight, but the unit toggle between metric and imperial gives him pause mid-flow. It's a small friction point, but for a user already uncertain about trusting the recommendation, doubt at the input stage compounds into doubt at the result.

The Brand-Loyal Returner

The Uncertain Measurer

Shops by brand, not by size

Wants help, gets stuck on inputs

Amaka shops by brand and thinks in brand-specific sizes. The brand-selection screen felt designed for her, but analytics showed she still dropped off there, because her preferred brands were often unlisted. She bore the cost of a screen that had no accuracy benefit.

Omar knows his height and weight, but the unit toggle between metric and imperial gives him pause mid-flow. It's a small friction point, but for a user already uncertain about trusting the recommendation, doubt at the input stage compounds into doubt at the result.

Three archetypes built from interview synthesis — each one maps to a specific failure point in the V4.2 flow

Where the journey broke down

Journey mapping the V4.2 flow made the problem visible in a way that analytics alone couldn't. We mapped the emotional arc across every screen; from the moment a shopper clicked "Find my size" through to the recommendation reveal, and marked where confidence dropped, where confusion spiked, and where users abandoned the flow entirely.

Two moments stood out. The brand selection screen caused a sharp confidence dip for users whose brand wasn't listed, and the recommendation reveal triggered disbelief for users whose result didn't match their expectation. Neither moment had any recovery mechanism in V4.2.

User journey map

V4.2, women's upper body

Eleven real screens rather than lifecycle stages, so the funnel data and the emotional arc can be read in the same place.

Screen Still in flow Confidence What happened, and how we know V5.0
00Find my size 100%opened the tool Curious Taps into the flow from the product page, expecting something quick. Analytics Out of scope
01Your measurements 67%−33 pts First doubt A third leave before entering a single measurement. An entry point and value proposition problem rather than a flow problem. Analytics Out of scope
02Belly shape 64%−3 pts Steady The illustrations tested well. The arrows and buttons around them read as clutter. Usability Arrows removed
03Hip shape 61%−3 pts Steady Same pattern as the screen before it, and the same fix. Usability Arrows removed
04Bra size 57%−4 pts Dipping A forty cell table, and female testers reported discomfort at being asked for this at all. Usability Help text added
05Your age 53%−4 pts Dipping Purpose unclear without the help text. "Why does it need this to sell me a sweatshirt?" Usability Help text kept
06Brand selection 36%−17 pts Sharp drop The steepest fall anywhere in the flow. Testers could not find their brand in the grid or in "More brands", and no recovery path existed. Analytics and usability Removed
07Brand summary 1 33%−3 pts Low Reads back a brand and a size the shopper never entered. Judged redundant. Usability Removed
08Brand size 25%−8 pts Low Brand data sits outside the body data mental model entirely, scoring zero against it. Card sort Removed
09Brand summary 2 20%−5 pts Low The same summary screen a second time, three questions from the end. Usability Removed
10Your best fit 14%−6 pts Splits in two Six more points lost at the last step, among shoppers who had answered every question. Of those who did see a result, the ones who got the size they expected were relieved and the ones who did not went to the size chart instead. Analytics and empathy mapping Reveal rebuilt

User journey map

V4.2, women's upper body

Twelve real screens rather than lifecycle stages, so the funnel data and the emotional arc can be read in the same place.

Screen Still in flow Confidence What happened, and how we know V5.0
00Find my size 100%opened the tool Curious Taps into the flow from the product page, expecting something quick. Analytics Out of scope
01Your measurements 67%−33 pts First Doubt A third leave before entering a single measurement. An entry point and value proposition problem rather than a flow problem. Analytics Out of scope
02Belly shape 64%−3 pts Steady The illustrations tested well. The arrows and buttons around them read as clutter. Usability Arrows removed
03Hip shape 61%−3 pts Steady Same pattern as the screen before it, and the same fix. Usability Arrows removed
04Bra size 57%−4 pts Dipping A forty cell table, and testers reported discomfort at being asked for this at all. Usability Help text added
05Your age 53%−4 pts Dipping Purpose unclear without the help text. "Why does it need this to sell me a sweatshirt?" Usability Help text kept
06Fit preference 49%−4 pts Dipping Seven options as radio buttons with no clear next step. 2 of 12 testers were unsure how to continue. Usability Buttons and CTA
07Brand selection 36%−13 pts Sharp Drop The steepest fall after the entry screen. Testers could not find their brand in the grid or in "More brands", and no recovery path existed. Analytics and usability Removed
08Brand summary 1 33%−3 pts Low Reads back a brand and a size the shopper never entered. Judged redundant. Usability Removed
09Brand size 25%−8 pts Low Brand data sits outside the body data mental model entirely, scoring zero against it. Card sort Removed
10Brand summary 2 20%−5 pts Low The same summary screen a second time, three questions from the end. Usability Removed
11Your best fit 14%−6 pts Splits in Two Six more points lost at the last step, among shoppers who had answered every question. Of those who did see a result, the ones who got the size they expected were relieved and the ones who did not went to the size chart instead. Analytics and empathy mapping Reveal rebuilt

Two points in the flow carried the losses. Brand selection cost 13 points in a single step and had no recovery path for shoppers whose brand was missing. The last step cost another six, among shoppers who had already given us everything we asked for.

Usability test findings mapped to each screen in the V4.2 flow

What the usability tests found, screen by screen

We ran moderated usability sessions with 12 participants across the full V4.2 flow, conducted remotely via Google Meet. The findings were specific enough to drive direct design decisions, not just general friction, but screen-by-screen evidence of where and why users lost confidence. The clearest was the progress indicator: 8 of 12 testers could not tell how far through the flow they were, which is why the ellipses came out in V5.0.

What empathy mapping revealed about trust

Empathy mapping sessions after the usability tests and interviews helped us understand what users were thinking and feeling at the moments that mattered most. The pattern was consistent: when a recommendation contradicted a user's expectation, they immediately looked for a reason and found nothing.

The product did explain itself, just not at the moment it mattered. There was a confidence figure and a line about what similar shoppers bought, but it sat under the recommendation in small text, it was the same sentence for everyone, and it never addressed the one thing the shopper had brought with her: the size she already believed she was. A generic explanation answers a question nobody asked. What they wanted to know was not "how does this work", it was "why is this different from what I expected".

Empathy map

The recommendation reveal

Scoped to one screen rather than the whole product, because the whole product mapped to roughly neutral and hid the moment that failed.

Says

  • "I know I'm usually a medium, but this thing told me large."
  • "The fit was surprisingly accurate for me."
  • "I'd want more detail on how it worked this out."
  • "I wish there were more options for different body types."

Thinks

  • How accurate is this, really?
  • Where is large coming from? Nothing here explains it.
  • Would I trust this enough to skip the size chart?
  • Has anyone my size actually reviewed this item?

Doesthe evidence

  • Opens the retailer's size chart in another tab
  • Scrolls back to the product page to cross-check measurements
  • Moves back and forth between the recommended size and the expected one
  • Adds the size they always buy, not the size they were given

Feels

  • Relieved when the result matches what they already believed
  • Sceptical when it does not, with no way to resolve it
  • Frustrated at answering ten questions for an answer they will verify anyway
  • Confident enough to buy only once the two agree

The DOES quadrant was the evidence. Shoppers who had just been given a size then went to the size chart, cross-checked the product page, and moved back and forth between options. People who believe an answer do not go looking for a second one.

Empathy mapping revealed that the trust gap wasn't about accuracy — users didn't distrust Fit Finder because it was wrong, they distrusted it because it didn't explain itself

I know I'm usually a medium, but this thing told me large. I didn't really trust it, so I just guessed anyway.

— Usability test participant, 2022

V4.2 funnel across the 20 retail partners later used in the A/B test. 100% is shoppers who opened Fit Finder. A third leave on the first measurement screen before entering anything. Of those who stay, the steepest single drop is Brand selection, and completion keeps falling across every subsequent brand screen.

What analytics and the similarity matrix confirmed

Analytics review revealed a clear pattern. Shoppers who made it through measurements, body shape and personal data abandoned at a disproportionate rate once they hit the brand screens. The steepest single drop after the entry screen was Fit preference at 49% to Brand selection at 36%, a 13 point fall in one step, and completion kept eroding across three more brand screens down to 20%.

It did not stop there. Only 14% ever reached a result: another six points lost at the final step, among shoppers who had already answered every question. Whether that was fatigue or the load time before the reveal, we never isolated it. What the qualitative work had already told us was that reaching a result and acting on one were different things. Two problems in one funnel, and they set the agenda between them: fewer screens to reach a recommendation, and a recommendation worth reaching.

Card sorting results: shoppers grouped height, weight, body shapes and bra size together as body data, kept age and gender as a separate pair, and placed brand preference in a category of its own.

Card sorting provided the mental model evidence to support removing it. Shoppers grouped height, weight, body shapes and bra size together as body data, kept age and gender as a separate pair, and placed brand selection in a category of its own. We were running three mental models through one linear flow.

The similarity matrix made that pattern quantitative. With 10 respondents, the blue clusters show near-perfect agreement; height, weight, chest shape, belly shape, hip shape, and bra size were grouped together by almost everyone. Brand information scored zero across those same groupings. The data wasn't ambiguous: brand data belonged to a mental model of its own, and the V4.2 flow had collapsed three into one.

Similarity matrix from the card sorting exercise — blue clusters show how users grouped personal measurements together and separated brand information into a distinct category

01

Brand data: no accuracy impact

The Data Science team's model analysis found brand selection made no measurable difference to accuracy. Four screens of user effort, no gain in the output.

02

Three mental models, one linear flow

Card sorting separated body data (measurements, shapes, bra size) from demographics (age, gender) from brand preference. V4.2 ran all three through one sequence as if they were the same question.

03

5 of 12 testers could not find their brand

Usability data showed 5 of 12 testers couldn't locate their brand in the selection screen. Either it wasn't listed, or they couldn't find it in "More Brands." Drop-off followed immediately.

04

4 of 12 testers ignored the recommendation

Post-flow interviews confirmed 4 of 12 testers didn't act on the size recommendation. They second-guessed it and checked the size chart manually. The problem wasn't the model, it was the interface.

Together these pointed one way. The brand screens were costing us a third of the funnel and returning nothing to the model.

Design

From sketches to a system

Six weeks of design work, moving from Crazy 8s through sketches, wireframes, and high-fidelity screens. The process ran alongside an accessibility audit and regular design critiques. I also involved junior designers throughout as a way to both pressure-test decisions and develop the team's skills in parallel.

1. Accessibility audit

1. Accessibility audit

Evaluated the existing V4.2 for heading hierarchy, focus order, screen reader compatibility, and WCAG colour contrast. Identified accessibility issues across the existing flow that needed resolution before any visual redesign.

2. Crazy 8s & sketching

2. Crazy 8s & sketching

Fast ideation on alternative approaches to the measurement input, recommendation reveal, and trust-building moments. Eight concepts in eight minutes, then down-selection based on feasibility and user insight alignment.

3. Wireframes & flow validation

3. Wireframes & flow validation

Mapped the simplified flow without brand screens. Validated with the PM and engineering lead before moving to high fidelity. No wasted polish on concepts that wouldn't ship.

4. High-fidelity design & ARES Design System

4. High-fidelity design & ARES Design System

Designed within, and contributed to, the ARES Shopping Suite design system. Built reusable components for measurement inputs, recommendation cards, and progress indicators that could scale across Snap's AR tools.

5. Prototype & usability testing

5. Prototype & usability testing

Interactive prototype tested before handoff. The reveal screen dropped the competing size bars for a single recommendation, and added fit context for garments that run large or small.

From research to decisions — annotated V5.0 screens

The V5.0 recommendation reveal screen addressed the trust gap directly. One size rather than two competing options, a plain-language line grounded in purchase behaviour at scale, and a fit note for garments that run large or small.

Every change in V5.0 was driven by a specific finding from usability testing; brand screens removed, radio buttons replaced with buttons, arrows and ellipses stripped from body shape screens, and help text added to the bra size screen to reduce discomfort

Fewer screens, and one answer

The problem statement had two halves, and each got one structural decision. Everything else in V5.0 was a change to a screen. These two changed what the product asked for and what it said back, so neither was mine to make alone.

1. Removing the brand screens

Four screens asked shoppers to pick a brand and their size in it, and cost 29 points of completion between them. I took the question to Data Science before any wireframing began: does this input earn its place in the model? It did not.

Data Science. Brand input made no measurable difference to recommendation accuracy, so removing it degrades nothing. It was a feature we had never been able to justify keeping.

User Research. 5 of 12 testers were unsure which brand to select in V4.2, and there was no way past the screen without guessing. Cutting it removes the confusion rather than explaining it.

Product. Four fewer screens is four fewer places to lose someone, but retailers will ask where brand selection went, so we need an answer ready before it ships.

Engineering. Those screens carried the oldest code in the flow. Removing them takes a maintenance burden with them.

The decision held. What changed was the rollout: the PM's point is why we prepared the client-facing reasoning ahead of the release rather than after it.

2. One recommendation, not two

V4.2 showed the recommended size beside an alternative with its own percentage. On the one screen built to remove doubt, that reads as the model hedging. This decision came later than the first, out of prototype testing rather than research, and it needed a different set of people.

Data Science. The model returns a distribution, not an answer. Showing one size is defensible where the top option clears a confidence threshold. Below that, the screen needs a different treatment.

User Research. 4 of 12 testers did not act on the V4.2 recommendation. Retested against a single size with no competing option, that dropped to 1 of 12.

Product. A single confident answer could lift add-to-cart and raise returns at the same time, which trades away the metric the product is sold on. Returns need segmenting in the test.

Client Solutions. Fit Finder sits on the retailer's own product page. Clients will notice the second size has gone and will want the reasoning, so it cannot ship as a silent change.

Engineering. The below-threshold behaviour has to be specified, or the component goes to production with an undefined state.

Top, the confident state. Bottom, the between-sizes state, where saying so plainly is a better answer than picking one and hoping.

We planned to show one size, always. Data Science pointed out that the model returns a probability for every size, so a single answer only holds when the top one is clearly ahead. What shipped was "show one size when the model is clear about it." A top size at 45% is a confident answer if the next one is 25%. At 55% it is a coin flip if the next one is 45%. They set where that line sat, not design.

What we owed the brand-loyal shopper

Removing the brand screens took away something one archetype had explicitly asked for. That deserves an answer rather than a shrug.

The screen was already failing her. Her brands were frequently unlisted, which is why she was dropping off at the step supposedly built for her. She was paying the full cost of a feature that worked for a fraction of the catalogue.

And we could deliver the benefit without the input. What she wanted was brand-specific calibration. What she was being asked for was labour. Garment-level fit context on the result screen, this style runs small, most shoppers size up, gives her the same answer, works for every brand rather than only listed ones, and costs her nothing.

Design system contribution

Fit Finder V5.0 was designed within the ARES Shopping Suite design system; a shared component library built to unify Snap's AR tools including virtual try-on and interactive product displays. I created the sizing-specific components (measurement inputs, recommendation cards, fit preference selectors) so they could be reused across future ARES products without redesign.

Fit Finder components built into the ARES Design System for reuse across Snap's Shopping Suite

Results

Three months, one decision

V5.0 launched in Q1 2023. Over a three-month A/B test across 20 retail partners, including global clients in apparel, the PM and Data Science team tracked completion, conversion, returns and satisfaction, with three of those set as rollout criteria.

Before the test began, the team aligned on three conditions for a full rollout decision: statistical significance at 95% confidence on conversion rate; no regression on return rate; and a neutral or positive shift in user satisfaction. All three were met, with return rate clearing its bar comfortably.

Completion rate

14% → 39%

14% → 39%

Conversion rate

+14%

+14%

Return rate

–12%

–12%

Each number alone could mislead. Completion alone means more people saw an answer, not that they believed it. A conversion uplift can mean people were persuaded into things that don't fit. A return rate drop can mean fewer people bought at all. Measured together across the same 20 partners, they describe one chain: more shoppers reached a recommendation, more of them acted on it, and fewer sent the item back. Completion was the figure the problem statement opened on.

Removing the brand screens was the largest single reduction in drop-off, because it deleted the steps where the losses were happening. The rebuilt reveal, with a plain explanation and a fit note for garments that run large or small, was flagged in qualitative feedback as a meaningful trust signal. Two decisions, one from research and one from prototype testing, pulling the same way.

V5.0 variant, women's upper body, across the 20 A/B partners. Six screens instead of ten.

With all three success criteria met and statistical significance confirmed, the PM and engineering teams moved to full rollout across all retail partners. V5.0 stayed the baseline for the rest of my time on the product, and the front door problem the funnel had been flagging since 2022 became the brief for the next cycle.

Reflection

With more time


The front door we never opened

The front door we never opened

A third of shoppers left on the very first screen, before entering a single measurement. It is the single largest loss anywhere in the funnel and we scoped it out, because it is an entry point and value proposition problem rather than a flow one. That was the right call for this release. It was the wrong thing to still be true two versions later.

What we removed and did not replace

What we removed and did not replace

Cutting the brand screens was right on the numbers, but one group had asked for brand-specific sizing and got nothing back. We argued garment-level fit context covered the same need without the work, and the aggregate supported it. We never segmented the test by shoppers who used to pick a brand. An aggregate win can hide a segment going backwards.

Accessibility auditing the old version, not the new one

Accessibility auditing the old version, not the new one

I audited V4.2 up front, which told us what was already broken. I never re-ran those checks on the new screens while they were still wireframes. Contrast and touch target issues surfaced at final review instead, so decisions had to be revisited late and small inconsistencies got into the system.

Isolating which change did the work

Isolating which change did the work

The test told us V5.0 beat V4.2 on every metric we tracked. It could not tell us how much came from the brand screens, how much from the reveal, and how much from the general cleanup, because we shipped them as one variant. I'd want an isolating cell: brand screens removed, nothing else.

Design system contribution scope

Design system contribution scope

Contributing to ARES while simultaneously shipping Fit Finder meant some components were built for the project first and generalised after. A few of them weren't quite right for other contexts and needed rework. A clearer split between project-specific and system-level work from the start would have saved that.

Fit Finder (V4.2), Upper body / Woman — October 2022

Role and collaborators

My role and the wider team

I was design lead on the project. That meant the accessibility audit, wireframes and high-fidelity design, the design system contribution, prototyping and developer handoff. On research I wrote the plan with our User Researcher and ran the synthesis with her; she facilitated the interviews and moderated the usability sessions. The decision to remove the brand screens was mine to push for, and I validated it with the PM and Data Science team before it touched a design file.

Engineering Lead + Developers

Owned implementation, managed the developer handoff process, and ran QA across web and mobile.

Data Scientists

Ran analytics review, validated the recommendation model changes, and led A/B test analysis.

User Researcher

Facilitated the interview programme, ran empathy mapping sessions, and moderated usability testing.

Product Manager

Scoped the project, aligned stakeholders across Snap and the retail clients, and oversaw the A/B test.