LLenny's Podcast
← All frameworks
InnovationGina Gotthilf (Latitud, Duolingo)

Minimum Viable Experiment + The Dogfood Gate

Ship the cheapest version of an experiment — but use it yourself before you trust its null result

Difficulty
Moderate
Time to result
~weeks to results
Steps
5
Confidence
94%

Growth teams rank experiments by expected return over time investment, which correctly kills big-ticket ideas — until a lean 'minimum viable experiment' version of a killed idea gets shipped so stripped-down that it fails to test the actual hypothesis, and the team wrongly writes the idea off. The fix is a dogfood gate: before an experiment can produce a verdict, someone on the team must personally experience it. A null result from an experiment nobody used is not evidence.

Origin

Gina Gotthilf's account of the Duolingo growth team's badges failure — the mistake she describes as her genuine one, in contrast to the interview-answer kind of mistake that flatters you.

Core principles

  • 01Rank experiments by expected user/DAU return against time investment — high-effort items lose to lower-hanging fruit.
  • 02A minimum viable experiment is legitimate, but only if it preserves the mechanism the hypothesis depends on.
  • 03A stripped MVE that removes the mechanism produces a false negative, not a cheap answer.
  • 04Dogfooding is the cheapest test that exists and it catches false negatives before they cost you a year.
  • 05Some wins are compounding platforms, not point features — killing them costs more than the experiment's own upside.

How to run it

  1. 1

    Rank every experiment by expected return over time investment

    Score each idea by how many users or DAU you expect it to move, divided by the engineering and design time it demands. Work the lower-hanging fruit first.

    Watch out This ranking systematically suppresses ideas whose payoff is a platform rather than a point win — you must manually re-examine those.

  2. 2

    Design the leanest version that still tests the hypothesis

    Before building the full feature, ask what the minimum viable experiment is. Then, critically, name the mechanism the hypothesis relies on and check the lean version still contains it.

    Pro tip Write the hypothesis as 'X works because of mechanism M.' If the MVE deletes M, the MVE is invalid.

    Watch out Duolingo's badge MVE deleted every mechanism badges depend on — pride in the achievement, a collection to complete, and social visibility.

  3. 3

    Dogfood the experiment before reading the result

    Someone on the team must actually use the experiment as a user before the result is accepted. Duolingo's growth team was disciplined about prioritization and write-ups but had never actually used its own experiments — one minute of dogfooding would have exposed the badge as lame.

    Pro tip Trust the flinch. If using it feels wrong or unexciting, that instinct is data, and you don't need three colleagues to validate it.

    Watch out Rigorous process (prioritization, write-ups, significance) can coexist with never touching the product — the process gives false confidence.

  4. 4

    Treat a null result from an untested experiment as unresolved, not closed

    If an experiment failed and nobody had used it, the idea is not dead — the implementation is. Re-open, rebuild with the mechanism intact, and re-run.

    Watch out Duolingo moved on for about eight more months without looking back — the cost of a false negative is measured in quarters, not sprints.

  5. 5

    Ask whether the win is a platform

    When a feature does land, check whether it creates new leverage. Badges weren't just a metric bump: once users want badges, you can attach any behaviour to them — invite friends, make purchases, anything.

In the wild

The badge nobody wanted

The growth team wanted badges — they were pervasive in every top game they studied. ROI ranking said the time sink was too high, so Gotthilf blocked the experiment for roughly six months. When they finally ran it, they ran the leanest possible version: you sign up, you get one badge, a girl with a balloon. It produced no results. Nobody is proud of signing up, there was no collection to build, and nothing to show anyone — every mechanism that makes badges work had been engineered out.

The team wrote badges off and moved on for another eight months. When they revisited and built badges properly, badges moved almost every company metric positively — including some they hadn't expected — and became a platform they could attach any user behaviour to.

Discovering they had never dogfooded

Reviewing the badge failure, the team realized they had never used their own experiments. Gotthilf, coming from marketing rather than product, had not even fully understood the term dogfooding.

Dogfooding became a standing part of the growth team's practice, and she still has to remind engineers at Latitud to do it — including ex-Nubank and ex-Uber PMs who all agree it's obvious and still forget.

Common mistakes

Letting the MVE strip out the mechanism

A cheap experiment that removes the thing the hypothesis actually depends on returns a false negative, and false negatives are more expensive than the experiment you avoided running.

Confusing process rigor with product contact

Excellent prioritization, hypothesis write-ups, and significance testing can all be in place while nobody on the team has ever touched the feature — which is precisely how the badge result went unchallenged.

Deferring to research instead of your own instinct

PMs often reach for user research when the faster signal is using the thing and noticing what feels wrong. Gotthilf's rule: when you don't get something, other people at scale probably don't get it either.

Never revisiting a killed experiment

'We tested it, it didn't work, we moved on' is only valid if the test was valid. Eight months of unquestioned closure was the real cost of the badge failure.

Is it for you?

Best for

Growth and product teams running a prioritized experiment backlog who keep getting flat results from lean tests

Not ideal for

Genuinely unbuildable or regulated experiments where a lean version cannot legally or technically exist

From the transcript

since we ranked all of our experiments in terms of Roi and return being like how many users we think we're going to get

37:30

there's MVPs there's like minimum viable experiments we don't have to run a whole badges thing

38:00

no one is proud of signing up it's not an exciting moment and you don't even have badges to collect

38:30

why are we not testing our experiments and so like that became part of our of our practice

39:30

when you don't get something probably other people don't get it too at scale

43:30

From the episode

Scaling Duolingo, embracing failure, and insight into Latin America’s tech scene

Gina Gotthilf (Latitud, Duolingo)