Case study, Solo Project, 2026

Roam: Trip planning by AI, grounded in reality

A live AI trip planner that checks every stop against real places and travel times.

Roam in 60 seconds
01

Problem

AI trip planners write itineraries that read well and fall apart: places in the wrong city, days you cannot actually walk. Anything can generate a plan. The real job is making sure the plan holds up.

02

What I did

Designed and built it alone in 48 hours, from tokens and components to the system prompt and a React front end built from the live Figma file. Then eight weeks of fixing what real use exposed.

03

Result

15 of 18 testers reached a saved plan unaided, in a median 3m 40s, rating ease 6.0/7. Testing showed trust held even when a plan was wrong, so every stop now has to be real and belong in the trip.

RoleSolo designer and builder
Timeline48 hours, then 8 weeks
ScopeResearch to a live MVP

Cost per plan

~€2.29

~€2.29

Time to a saved plan

3m 40s

3m 40s

Testers who finished

15/18

15/18

Commits since launch

250+

250+

Cost per plan measured from the server log of a real 3-day trip, Google APIs at list price. Testing: unmoderated remote study, 18 participants, directional.

The product

One designer. No engineers. Shipped.

Enter a destination, set your dates, travellers and budget, and Roam generates a real day by day itinerary with live places, real travel times between stops, and an estimated nightly price range for where you stay. Roam is mobile-first by design, built for the phone people actually plan trips on.

Claude read the live Figma spec through MCP and built the React front end from it; the tools, the phases and the cost model are all below.

This is the live app, not a prototype. The core flow, from brief to saved plan, is complete; everything outside it is a placeholder for now. →

Context

A trip to Albania that became a product question

Group trips still get planned across a WhatsApp thread, a Google Doc, a few blogs and a spreadsheet. Most travel apps search for things you already know you want. None of them help you work out what you want in the first place.

After planning a trip to Ksamil in Albania as a group of four, juggling ferry times, stays, beaches, restaurants and daily budgets across a handful of tools, none of which helped four people decide what the days should actually be, the question felt real: could a single AI-native interface replace all of that without feeling like yet another chatbot?

Roam was the attempt to find out, built from scratch with a hard budget of 48 hours.

Context

A trip to Albania that became a product question

Group trips still get planned across a WhatsApp thread, a Google Doc, a few blogs and a spreadsheet. Most travel apps search for things you already know you want. None of them help you work out what you want in the first place.

After planning a trip to Ksamil in Albania as a group of four, juggling ferry times, stays, beaches, restaurants and daily budgets across a handful of tools, none of which helped four people decide what the days should actually be, the question felt real: could a single AI-native interface replace all of that without feeling like yet another chatbot?

Roam was the attempt to find out, built from scratch with a hard budget of 48 hours.

Discovery

Planning a trip is scattered by default

Before designing anything, I planned the Albania trip and paid attention to where it fell apart. Then I checked it against national surveys in the US and Germany.

US, 2025: AI sits near the bottom but grew fastest, up 30% in a year. Amadeus, 2,000 travellers.

Germany, 2024: portals lead, a third still visit a travel agency, AI is at 3%. Bitkom Research, 954 holidaymakers.

These surveys only measure where a trip starts, though. What happens next, when a scattered set of ideas has to become a day you can actually walk, is where my own planning came apart.

AI is on both lists, and it does not survive a real day

In the same Amadeus survey, 27% said AI had returned inaccurate information, and Amadeus's own President of Hospitality said gen AI tools are not yet ready for trip planning prime time. When InsureMyTrip tested ChatGPT, Gemini and Google's AI on a real week-long trip, the plans recommended a restaurant that does not exist and laid out days with travel times nobody could actually make.

Independent testing

"The itineraries often looked polished and logical on the surface... closer inspection revealed that both big and small details were often inaccurate."

Sara Boisvert, InsureMyTrip, reported in Forbes, 2026

That became the thesis for the whole product. Anything can generate a plan. The design problem is guaranteeing the plan is real: places that exist, prices that are close, and a day whose stops you can actually get between before dinner. Trust, not text, is the job. Four failures showed up again and again, and each one became something Roam had to answer.

Problem

The four failures Roam had to answer

01

Trip planning has no single home

A dozen tabs: a group chat, a doc nobody updates, blogs, a budget sheet. Roam turns one short brief into a full day by day plan with places, costs and a map.

02

AI bolted onto the wrong moment

Booking tools assume you already know where you are going, then bolt AI on after. Roam is AI-first from the first screen, built for "help us decide where to go."

03

Plans that look right but fall apart

Ask any chatbot and it reads well, then breaks: lunch across town, stops with no gap between them. Roam checks each plan against real travel and fixed meal times.

04

Change one stop, the plan breaks

Swap a museum for a market in most tools and the timings fall out of sync. In Roam, change any stop and the schedule recomputes itself around it automatically.

Tech stack

Six platforms. One pipeline.

Six platforms, each feeding the next. Claude read the design straight from Figma, so design and code never separated. The app is React 19 and TypeScript, bundled with Vite.

Claude

Cowork

API

Code

Ideation

Itinerary & pricing

Component build

Figma

Make

Design

Agent

Token Studio

MCP

Wireframes

Design system

Design assistance

Design tokens

Design bridge

Google APIs

Places

Routes

Static Maps

Location data

Travel times

Map view

Cursor

Agent

Editor

Front-end build

Edge-case fixes

Vercel

Deploy

Functions

Edge CDN

Production hosting

Serverless API

Photo caching

UXtweak

Study

Recordings

Unmoderated tests

Screen replay

Claude

Cowork

API

Code

Ideation

Itinerary & pricing

Component build

Figma

Make

Design

Agent

Token Studio

MCP

Wireframes

Design system

Design assistance

Design tokens

Design bridge

Google APIs

Places

Routes

Static Maps

Location data

Travel times

Map view

Cursor

Agent

Editor

Front-end build

Edge-case fixes

Vercel

Deploy

Functions

Edge CDN

KV

Production hosting

Serverless API

Photo caching

Place cache

UXtweak

Study

Recordings

Unmoderated tests

Screen replay

Process

48 hours, phase by phase

Seven phases, each leaning on a different tool. When a phase ran long, scope was cut, not quality.

Hours Tool Phase What happened
0–2 Claude Ideation Stress-tested the concept in a Claude session. The output was a set of decisions: three surfaces, two integrations, one hard time limit.
2–4 Figma Make Wireframes Nine screens from a plain-language brief. Several survived almost unchanged, which moved the real work to flow rather than layout.
4–20 Figma Agent Hi-fi design Agent populated the frames with real content, so the design went from placeholder boxes to something that felt like a product.
20–24 Token Studio Design tokens to CSS The whole visual system exported as CSS custom properties. Change a colour in Figma, it changes in the app.
24–28 Anthropic API
Google Places
Planning logic and prompt Wired up the planning layer and wrote the system prompt, where the two-phase "ask, then generate" model was defined. Google Places sat behind it so every suggestion resolved to a real location.
28–40 Figma MCP Design to code Claude built each component straight from the live spec, read through MCP. This is the phase that made handoff feel solved rather than tedious.
40–48 Vercel Build and deploy Tight iteration on edge cases and loading states, then a production deploy in under two minutes.

Wireframes

The structure, decided in wireframes

Hours two to four went into wireframes, not visuals. Figma Make turned a plain-language brief, a travel app with a trip form, an AI planning flow and an itinerary view, into nine rough screens. Grey boxes, real structure, enough to react to.

Click through the wireframes Figma Make generated from a plain-language brief. Several survived, almost untouched, into the final build.

Design foundation

A system before the screens

Before any screen went to high fidelity, I built the foundation everything else would inherit from: a set of tokens and a component library.

Design Tokens

Colour primitives were aliased to semantic tokens, each graded for AA or AAA contrast, so no component ever touched a raw hex. Plugged into Token Studio in Figma, the tokens export to CSS custom properties, so one colour change flows through the whole file and into the app.

Design System

On top of the tokens sat the component library: buttons, tags, cards, inputs, confidence badges and meal pills, each designed once. Figma Agent generated the variant matrix, leaving me the decisions that needed judgement. Every screen drew from that set, so one fix landed everywhere.

The token reference on the left, and the component library it feeds on the right, from colour primitives through to finished buttons, cards and badges.

The token reference on top, and the component library it feeds below it, from colour primitives through to finished buttons, cards and badges.

Tokens and components, straight to code

The same pipeline, traced through two components. The meal tag follows a colour, the trending card follows shape and elevation. Each time, the value set in the design system is the exact token the shipped CSS renders from.

Figma · Design System
Dinner
MealTag · Property 1 = Dinner · selected
Fill #FFF1F2
Stroke #FDA4AF
Text #9F1239
Radius 9999

The Dinner meal tag, selected. Its colours live in the design system, not on the component.

tokens.css
/* Meal */
--meal-dinner-fg: #9F1239;
--meal-dinner-bg: #FFF1F2;
--meal-dinner-border: #FDA4AF;

Each colour is one named token, the only place the raw hex is written.

components.css
.meal-tag--dinner {
color: var(--meal-dinner-fg);
background: var(--meal-dinner-bg);
border-color: var(--meal-dinner-border);
}

The component asks for the token by name. It never touches a raw colour.

Figma · Design System
Santorini, Greece
Whitewashed clifftop villages overlooking the Aegean
TrendingCard · selected
Fill #FFFFFF
Radius 16
Elevation sm

The trending card, selected. Its shape, surface and shadow all come from the system, not hard-coded values.

tokens.css
/* Elevation & shape */
--radius-lg: 16px;
--surface-default: #FFFFFF;
--shadow-sm:
0 1px 3px 0 rgba(15,23,42,0.08),
0 1px 2px 0 rgba(15,23,42,0.04);

Shape and elevation are tokens too. The shadow is one compound token carrying a whole two-layer elevation.

components.css
.trending-card {
border-radius: var(--radius-lg);
background: var(--surface-default);
box-shadow: var(--shadow-sm);
}

The card sets no pixels of its own. It reads its radius, surface and shadow from named tokens.

Design

Three acts, and the prompt behind them

Roam's interface follows a three-act structure: tell it about your trip, let it plan, then explore and refine. The brief opens with a structured form rather than a blank chat box, because most people do not want to type an essay to start a trip, and the form doubles as the model's context, so it can generate straight away. The interest chips are generated for the destination you type, so Marrakech offers souks and Bergen offers fjords, and the first set a place gets is the set everyone sees from then on.

Work in progress. Early passes at each surface, where the structure and flow got settled before the design system brought everything to final fidelity.

Figma's agent doing the repetitive build. A plain-language brief, and it duplicated and edited three detail screens in parallel while I stayed on direction.

Early passes at each surface, before the design system.

Figma's agent building three detail screens in parallel from one brief.

The shipped app, from a short brief to a saved plan.

Hometrending + saved trips Plan a tripdestination, dates, budget, interests Generatingsearching real options Accommodationhotel by price tier Generatingbuilding both plans Comparisontwo plans, pick one Map viewstops on a static map Finalise & savereservations, export, save AI + Google Places verification Every stop is checked against Google Places. Only real, correctly-located places survive, then real travel times are added between them. Swap placeslist of swappable stops Swap placepick a real alternative saved trip
Hometrending + saved trips Plan a tripdestination, dates, budget, interests Generatingsearching real options Accommodationhotel by price tier Generatingbuilding both plans Comparisontwo plans, pick one Map viewstops on a static map Finalise & savereservations, export, save AI + Google Places verification Every stop is checked against Google Places. Only real, correctly-located places survive, then real travel times are added between them. Swap placeslist of swappable stops Swap placepick a real alternative saved trip
Home
trending + saved trips
Plan a trip
destination, dates, budget, interests
Generating
searching real options
Accommodation
hotel by price tier
Generating
building both plans
AI + Google Places verification
Every stop is checked against Google Places. Only real, correctly-located places survive, then real travel times are added between them.
Comparison
two plans, pick one
Peel off to swap a stop: Swap places, pick a real alternative, then back to the comparison.
Map view
stops on a static map
Finalise & save
reservations, export, save
A saved trip loops back to Home.

The end-to-end flow. A short brief becomes a plan, then you compare the two pacing options, open the map, or peel off to swap a stop, and any swap recomputes the itinerary around it.

Designing the prompt as much as the interface

The product's behaviour was designed in the system prompt, not in a Figma frame. Getting itineraries that felt specific and appropriately scoped took as much iteration as any screen. Four decisions changed the output most.

01

Structured output

The model returns a consistent shape: time, place, duration, reasoning note. The UI renders it predictably regardless of destination.

02

Brief as context

Every field the user fills becomes a constraint the model reasons within, rather than a question it has to ask.

03

Budget as a constraint

Economy, Standard and Luxury are injected as constraints, calibrating stays, dining and activities to the same tier.

04

Reasoning first

A one-line rationale for every stop, so the output feels advisory, not generated. It matched the core insight from my Muse project: unexplained recommendations do not build trust.

After the sprint

What only real use reveals

Roam wasn't built in a day. It was built in two, then rebuilt over eight weeks and 250+ commits. That's where it got good.

The Vercel deployment log, one day of it. Each row is a commit that went straight to production, with a title that says what actually broke.

The Vercel deployment log, one day of it. Each row is a commit that went straight to production, with a title that says what actually broke.

Cutting the cloud bill

About a week in, Google Cloud credits were draining far faster than the traffic justified: €134 of €263 in seven days. Almost all of it was the place search, where every lookup asked for Enterprise-tier fields. So I cut them. Then I put two back.

Review count is the only thing separating a famous restaurant from a chain branch, and opening hours the only way to know a place is open when you arrive. Without them, dinner came from the same yakiniku chain twice and the app scheduled a government building at 22:15. Right on the numbers, wrong on the product. The saving that survived is caching, so a place is never paid for twice, and static photos, so the app only costs money when someone plans a trip.

The bill, line by line. Both Text Search tiers on one account: the cut to Pro, and the Enterprise calls I bought back. The free trial absorbed that month's total; it has since run out.

The bill, line by line. Both Text Search tiers on one account: the cut to Pro, and the Enterprise calls I bought back. The free trial absorbed that month's total; it has since run out.

Nothing cached yet ~€2.29
Place lookups62 calls · Enterprise
€1.98
Hotel search3 calls · Enterprise
€0.10
Photos~32 loaded
€0.21

Verifying every stop is what costs money, and both pacing plans are verified, not just the one you pick.

Mostly cached ~€1.12
Place lookupspartly cached, new places still bill
~€1.12
Hotel searchKV cache
€0.00
PhotosVercel CDN
€0.00

A place already looked up costs nothing. Different dates or interests still turn up new ones.

One full 3-day generation at Google's per-call prices above the free tier. Same tier on both sides; the saving is a shared cache. Both figures are read from the server log of a real Valencia trip.

When it runs out, and when it breaks

Cheap multiplied by unlimited is still unlimited, so Roam now plans a set number of trips a day. At the limit it offers a complete example trip instead of an error, because the product is busy, not broken. A check after every deploy tells me when it genuinely is broken.

The itinerary rulebook

The biggest body of work was a set of rules the model has to obey before a plan reaches the screen, enforced in the prompt and again on the server.

Rule Why it exists
No idle gaps The only space between two stops is the travel between them, checked to the minute by a script that replays a real saved plan.
Fixed meal times Breakfast at 09:00, lunch at 13:30, dinner at 20:00, the same on every day. The stops flex around them; the meals do not.
Opening hours Every stop is checked against its real opening hours for the whole visit. A place that would be shut is replaced; one that would close early is shortened.
No place twice Dedupe on the resolved place and on the brand, not the name. This caught a Lahore plan listing the fort and a palace inside it twice.
Distance-first routing Keep every stop within 15 km of your stay and reorder any day that doubles back, with Slow carrying fewer stops than Packed.
Balanced interests One stop per interest a day, and a cap per trip that scales with its length, so a plan never has three casinos or two shrines.
Resolve or drop A stop that fails verification is dropped rather than shown with a time and a travel leg, looking like one that checked out.
No numbers in prose No travel times, distances or visit lengths in the text. Every number on a card is measured, so the prose cannot contradict it.

When a correct rule still ships a wrong plan

A real place wearing another place's description. When a stop failed verification, the substitution replaced its name, address, photo, location, hours and category, and left the original description untouched. A members' club in Roppongi shipped described as a 24-hour ramen chain. Real place, real photo, fluent prose about a different business. Not hallucinated, mismatched, and it read perfectly.

A rule that could not fire. The check for stops scheduled at an hour they cannot keep was correct, and ran before two later passes that moved stops. Anything those passes touched shipped unexamined. A shrine went out at 21:10 and a design centre at 21:00, both hours after closing, under a rule written to prevent exactly that. It happened again one level up, in the pass that straightens a day. The second time I understood it was a shape, not an incident.

Each of those cost ~€2.29 to reproduce, because finding out meant generating a real plan against real APIs. So I built an offline harness that replays the actual scheduling code against a saved generation and reports meal times, route doubling-back, opening-hours violations and continuity to the minute. A change that used to cost a euro and a coffee now costs a second.

Reworked and fixed

Two things I got wrong early, both fixed. Unglamorous work, but it is the difference between a demo and something people can actually rely on.

Area What changed
Budget Moved the budget selector to the start of the flow. Set late, it meant regenerating the whole plan; upfront, it becomes a constraint the model plans within from the first call.
The map Shipped completely broken. Fixing it took a solo pass through Google Cloud: enabling Static Maps separately, attaching billing, and realigning a key that no longer matched the console.

User testing

What fifteen travellers taught me

The sprint proved I could ship it. Testing was where I found out whether anyone would actually trust it. Eighteen people set out to plan a trip to a city of their choice, on their own phones, unguided, while the screen recorded. Fifteen got to a saved plan, and most of what they planned held. When it went wrong, most people didn't notice, and that taught me the most.

Method

Unmoderated remote usability test on real phones, via UXtweak with screen recording. Success and time on task judged from the recordings, not a completion flag.

Task

One end to end task: plan a two-night trip to a city of their choice, choose where to stay, reach a finished plan and save it. Questions before and after.

Participants

15 people who plan their own leisure trips, a convenience sample. 18 started; the 3 who didn't finish aren't counted in the findings.

Measures

Completion, Single Ease, a trust rating, UMUX-Lite, would-follow and would-use intent, plus open text coded into themes. Small n, read as directional.

Fifteen people planned a full trip on Roam and saved it, across 13 cities on five continents, without needing help from me to get there.

15/18

15/18

reached a finished, saved plan unaided, recordings confirmed

6.0/7

6.0/7

mean ease of planning the trip (range 4–7)

5.8/7

5.8/7

mean trust that places are correctly located (range 4–7)

3m 40s

3m 40s

median time from a blank screen to a saved plan

What worked

01

The core promise landed, even for sceptics

All 15 reached a plan in minutes, and most called it easy. "It's convenient to see a skeleton of a trip in just a few questions." Even the participant who dislikes itinerary apps praised "the variety of locations and activities."

02

Useful pacing plans, easily missed

Eight of fifteen noticed them, usefulness 6.6 out of 7, and one wanted a third. From the video recordings, five of those eight actually toggled between them, so most people who found the choice went on to use it.

03

People stopped to read the reasoning

On 13 of 15 recordings, participants paused on the one-line rationale for each place. The reasoning-first decision paid off: the "why this" note gets read.

04

Seeing the map made it credible

Asked why the plan felt right, one participant just said: "A map with the itinerary is shown." The map itself did the convincing, and that is the thread the next finding pulls on.

Findings at a glance

Rated on Nielsen's severity scale, a blend of how bad, how frequent and how persistent. Each finding maps to one change.

What I found Severity The change it points to
People want a visible sense of control, not less automation.Raised by 9 of 15 Critical Surface the levers Roam already has: budget, timing and choosing between options.
The two pacing plans are useful but easy to miss.7 of 15 never noticed them Major Make the two plans easier to spot, and the Packed versus Slow difference visible at a glance.
Swapping was not recognised as the way to change the plan.4 of 15 missed it Minor Move the swap onto each stop so “change this” reads as an action.
Stops resolve to the wrong location. Three plans routed a day to another continent.3 of 15 plans · only 1 caught Minor Reject any stop outside the destination or more than 15 km from your accommodation, and surface the Google Places check so correctness is shown, not assumed.
Accommodation cards select but do not open to detail.3 tried it, 2 asked for it Minor Make cards expand to a detail view with photos and a price range.

The most revealing finding: trust held even when the plan didn't

Three of fifteen plans placed a stop on the wrong continent. Left: a first-timer's Day 1 ran from Lisbon to the US east coast; she never noticed and scored the plan 7 out of 7. Right: another participant's Day 3 sent him to the Azores, about 1,360 km into the Atlantic; he caught it at once and his trust fell to 4. The same bug, opposite reactions, and only the sceptic was protected by his own doubt.

What testing caught

"On day three the itinerary suggested an activity far outside Lisbon. The weird location was suspicious."

Participant 4, unprompted

Why it happened

The insight is sharper than "the AI hallucinated," because it didn't. Both places are genuine Google Places listings. The failure was resolution, not invention: a name matched a real entry in the wrong location, and the distance rule meant to keep a day inside one city let a 1,360 km outlier through. "Verify the place is real" and "verify the place belongs in this trip" turned out to be two different guarantees, and Roam only had the first.

The more uncomfortable half is the reaction. People never see the verification, so a confident tone carries the plan. One participant read an itinerary that would have flown her across the Atlantic, found it reasonable, and rated it a perfect seven. A high trust score partly measures how well a plan performs certainty, not how correct it is, which is exactly the risk the product exists to remove.

The most revealing finding: trust held even when the plan didn't

Three of fifteen plans placed a stop on the wrong continent. Left: a first-timer's Day 1 ran from Lisbon to the US east coast; she never noticed and scored the plan 7 out of 7. Right: another participant's Day 3 sent him to the Azores, about 1,360 km into the Atlantic; he caught it at once and his trust fell to 4. The same bug, opposite reactions, and only the sceptic was protected by his own doubt.

What testing caught

"On day three the itinerary suggested an activity far outside Lisbon. The weird location was suspicious."

Participant 4, unprompted

Why it happened

The insight is sharper than "the AI hallucinated," because it didn't. Both places are genuine Google Places listings. The failure was resolution, not invention: a name matched a real entry in the wrong location, and the distance rule meant to keep a day inside one city let a 1,360 km outlier through. "Verify the place is real" and "verify the place belongs in this trip" turned out to be two different guarantees, and Roam only had the first.

The more uncomfortable half is the reaction. People never see the verification, so a confident tone carries the plan. One participant read an itinerary that would have flown her across the Atlantic, found it reasonable, and rated it a perfect seven. A high trust score partly measures how well a plan performs certainty, not how correct it is, which is exactly the risk the product exists to remove.

That word, "feeling", is the point. The people who loved that Roam did everything for them and the people who wanted to steer are not actually in conflict. The automation should stay; what is missing is a visible sense of agency inside it. Tellingly, the participant who felt most in control praised "the high level of customisation given the few preferences I was asked to fill out," the same lightweight brief that left others feeling steered. The levers already exist. They are just not surfaced.

The takeaway

Fifteen sessions turned a shipped app into a ranked to-do list. Roam is fast, easy and persuasive, sometimes more persuasive than it should be. The work ahead is narrow and clear: give people the felt sense of control the automation quietly takes away, and make correctness visible, so a wrong stop cannot ride on a confident tone.

What users wanted

"Give the user more feeling of control. I understand the idea of the app is to let the machine do the work, but I would like to have the feeling at least."

Participant 5, on what he'd change

The theme that ran through everything: give people the feeling of control

Nine of the fifteen wanted more say over the plan, even while liking it. It would have been easy to read that as "add more settings." One participant reframed the whole finding.

Closing the second guarantee

The fix goes straight at the gap the testing exposed. Roam already checked that a place was real, and now it also checks that a place belongs in the trip. When it looks up a stop, a name match is no longer trusted on its own. The place is only accepted if it sits within 15 km of where you are staying, so the right name in the wrong part of the world is thrown out before it reaches the plan.

The 1,360 km outlier that slipped past testing is exactly the kind of thing it now catches.

That fix exposed a second one. The map refused to plot a stop it could not locate, but the itinerary list did not check at all, so an unresolved stop still looked like a verified one. A two-day Valencia plan shipped four of them.

Now a stop that fails the check is removed from the itinerary completely, and the day rebuilds around the gap.

Closing the second guarantee

Roam already checked that a place was real. Now it also checks that it belongs in the trip: a stop is only accepted within 15 km of where you are staying, so the 1,360 km outlier from testing is thrown out. That exposed a second gap. The map refused to plot an unresolved stop, but the list still showed it, and a two-day Valencia plan shipped four. Now a stop that fails the check is removed and the day rebuilds around it.

The theme that ran through everything: give people the feeling of control

Nine of the fifteen wanted more say over the plan, even while liking it. It would have been easy to read that as "add more settings." One participant reframed the whole finding.

What users wanted

"Give the user more feeling of control. I understand the idea of the app is to let the machine do the work, but I would like to have the feeling at least."

Participant 5, on what he'd change

That word, "feeling", is the point. The people who loved that Roam did everything for them and the people who wanted to steer are not actually in conflict. The automation should stay; what is missing is a visible sense of agency inside it. Tellingly, the participant who felt most in control praised "the high level of customisation given the few preferences I was asked to fill out," the same lightweight brief that left others feeling steered. The levers already exist. They are just not surfaced.

The takeaway

Fifteen sessions turned a shipped app into a ranked to-do list. Roam is fast, easy and persuasive, sometimes more persuasive than it should be. The work ahead is narrow and clear: give people the felt sense of control the automation quietly takes away, and make correctness visible, so a wrong stop cannot ride on a confident tone.

Reflection

What building and testing this taught me

Product behaviour is designed in the prompt

The decisions that changed output quality most were about how the model was instructed to behave, not layout or colour. Writing the prompt with the rigour of UX copy mattered more than any single screen.

Every number on screen is measured, not generated

Travel times come from real routes, opening hours from real listings, and the prose never contains a number it can't back up. It sounds like a technical rule. It is a trust decision, because a plausible number a reader can't check is exactly how a confident plan goes wrong.

Show the working

Roam verifies every place, but users never see that, so they trust the tone instead, and one trusted a plan that would have crossed an ocean. The verification exists. The first piece of it is now on every card, the Google rating and review count the check produced. The rest of it is the next job.

Automation and control are not opposites

Testing pushed me off my own assumption. People want the machine to do the work and to feel in charge of it. The interesting problem is agency inside automation, not one or the other.

Next steps

What comes next

The first job is the one testing ranked for me: make the two pacing plans and the swap on each stop easy to find, surface the levers people asked for, and show the rest of the verification rather than asking people to take it on trust. After that, the path is a real product, not a tidier prototype. That means a proper backend and accounts so trips persist, real booking so a saved plan becomes a trip you can take, and deeper coordination for a group planning together. That last one is where Roam started: four people, a ferry timetable and a trip to Ksamil.

A plan has to be real, it has to belong in the trip, and the people following it have to feel in charge of it. The clearest lesson came from one participant who grew suspicious of a stop far outside Lisbon. Roam's job now is to make that suspicion unnecessary.

Tech stack

Six platforms. One pipeline.

The 48-hour budget was not just a timeline, it was a design constraint: every decision had to answer "does this justify the time?" Six platforms covered the full stack, each feeding the next. Design and engineering never really separated, because Claude read the node data straight from Figma, so spacing, component properties and type translated into code without a manual handoff. The code itself is a React 19 app in TypeScript, bundled with Vite.

A live AI trip planner that checks every stop against real places and travel times.

Roam in 60 seconds
01

Problem

AI trip planners write itineraries that read well and fall apart: places in the wrong city, days you cannot actually walk. Anything can generate a plan. The real job is making sure the plan holds up.

02

What I did

Designed and built it alone in 48 hours, from tokens and components to the system prompt and a React front end built from the live Figma file. Then eight weeks of fixing what real use exposed.

03

Result

15 of 18 testers reached a saved plan unaided, in a median 3m 40s, rating ease 6.0/7. Testing showed trust held even when a plan was wrong, so every stop now has to be real and belong in the trip.

RoleSolo designer and builder
Timeline48 hours, then 8 weeks
ScopeResearch to a live MVP