LLenny's Podcast
← All frameworks
LeadershipKatie Dill (Stripe, Airbnb, Lyft)

Walk the Store: The Essential Journeys Quality Program

Name your 15 critical user journeys, give each a cross-functional owner, and score them on a fixed cadence.

Difficulty
Advanced
Time to result
~months to results
Steps
7
Confidence
95%

Products ship at their peak quality and then quietly regress as teams optimize their own slices in isolation. Stripe counters this by naming a finite set of critical end-to-end user journeys, assigning each an engineering + product + design owner trio, and having those owners periodically 'walk the store' — traversing the whole journey as a user would (search → website → docs → dashboard), friction-logging what they hit, scoring it against a rubric, and defending that score in a cross-functional calibration meeting.

Origin

Developed by Katie Dill at Stripe with CTO David Singleton, drawing on Dill's earlier work at Airbnb and Lyft. The friction-log component is Stripe's own practice, popularized by David Singleton. The 'walk the store' metaphor borrows from retail floor-walking; the scoring calibration is explicitly modeled on how managers calibrate performance reviews.

Core principles

  • 01Quality regresses in the wild — shipping is the high-water mark unless someone re-inspects.
  • 02Users experience journeys, not features; org boundaries are invisible to them.
  • 03Quality is a group effort — you cannot outsource it to QA or to one talented hire.
  • 04Firsthand pain is more visceral than reported pain; this augments user research, it does not replace it.
  • 05Judgment is what you hired for — the score can be qualitative without being worthless.

How to run it

  1. 1

    Pick a finite set of essential journeys

    Choose a tractable number of the user journeys that matter most (Stripe started with 15). It is deliberately not comprehensive — it is the set you can actually keep track of and hold people accountable for.

    Pro tip The number is arbitrary on purpose. Optimize for a list leaders can hold in their heads, not for coverage.

    Watch out Do not scope journeys to a single team's surface — a journey that fits neatly inside one org chart box is not a journey.

  2. 2

    Assign a cross-functional owner trio per journey

    Each journey gets an engineering, product, and design leader jointly responsible for its quality. Ownership is of the whole journey, not of the components each function happens to build.

    Pro tip Do the walkthrough together, not separately — one person notices load time, another notices copy inconsistency, another notices the off-design-system component.

  3. 3

    Walk the store end to end

    Owners traverse the journey as a user would, starting where the user actually starts — an internet search — through the marketing site, docs, and dashboard. They experience it, they do not review a deck about it.

    Pro tip Leaders should also do unscheduled ad-hoc walks of random flows outside the formal program; Dill and CTO David Singleton do these together, with him inspecting code and her inspecting experience.

    Watch out This is an additional input, not a substitute for user research or data.

  4. 4

    Friction-log against a template

    Fill a standard friction log: screenshots plus what you experienced, with each moment tagged for severity — from 'that was a nice touch' to 'consider a fix' to 'P0, fix now'. File bugs and route them to the owning teams as you go.

    Pro tip A shared template makes logs comparable across journeys, which is what makes calibration possible.

  5. 5

    Score with a qualitative rubric

    Roll the tagged moments into a summary score using a rubric covering usability, utility, desirability, and whether it reaches surprisingly-great. Stripe uses a colour system (e.g. yellow, yellow-green) rather than numbers or letters.

    Pro tip Colours stop people arguing about whether it is a six or a seven. The score is explicitly qualitative judgment, not a quantitative measurement.

    Watch out Numeric scales invite false precision and make people 'tied around the axle' on measurement instead of on the product.

  6. 6

    Calibrate scores in a Product Quality Review

    Owners present their walkthrough and score in a cross-functional PQR attended by design, engineering, and product leadership. The room debates whether the score is right — 'that felt worse than you described' or 'actually that hits the mark' — exactly as managers calibrate performance ratings.

    Pro tip Keep the room small enough to have a real discussion but broad enough that product marketing, engineering, and product are all represented.

    Watch out Without calibration, each team quietly grades on its own curve and the company-wide quality bar never converges.

  7. 7

    Publish scores and run it on a cadence

    Update a shared scorecard dashboard on a fixed cadence — quarterly at Stripe — so scores can be tracked over time. The cadence is the floor, not the ceiling; the hope is teams walk their journeys weekly on their own.

    Pro tip Quarterly is long enough for material change to appear between reviews and short enough to catch regressions.

In the wild

Upstream SEO fix discovered by walking the journey

Because the walk starts at an internet search rather than inside the product, a team discovered that their SEO and product framing did not match how they wanted users to understand the product later in the journey — a disconnect invisible from inside the dashboard.

The team fixed the upstream articulation, improving downstream outcomes later in the journey.

Skeptical engineers become quality converts

People from non-design functions who initially saw the program as unnecessary — focused on the technical execution of their piece — went through the walkthrough and saw the product from the user's lens for the first time.

They became advocates for the program, spreading the shared quality bar exactly as the 'quality is a group effort' principle predicts.

Common mistakes

Delegating quality to one function

Believing you can hire one incredibly talented person, or lean on QA, to own quality while everyone else carries on as before. Dill: you are sunk. Quality only holds if engineering, product, design, and marketing all inspect and care.

Reviewing features instead of journeys

Users almost never experience one surface in isolation. They learn about it, get to know it, decide to use it, then need it in a new way. Reviewing your own slice tells you nothing about the composite experience.

Turning the score into a metric war

Debating whether something is a 6 or a 7 diverts energy from the product. The score is a conversation starter for judgment, not an objective measurement.

Is it for you?

Best for

Product, design, and engineering leaders at a scaling company (100+ people) where teams have specialized into focused business areas and the end-to-end experience is degrading between org boundaries.

Not ideal for

Very small teams where everyone already uses the whole product daily, or pre-PMF products where the journey itself is still changing weekly.

From the transcript

they review these Journeys what we call walk the store

35:00

they friction log what they experience which I know David Singleton talked about on your podcast

35:30

we will come together in what we call pqr product quity review

43:00

we have landed on a color system

45:30

it's not meant to be an objective qualit quantitative score it is qualitative it is Judgment we hire people for judgment

46:00

From the episode

Building beautiful products with Stripe’s Head of Design

Katie Dill (Stripe, Airbnb, Lyft)