LLenny's Podcast
← All frameworks
Innovation

The Minimum Viable Experiment Trap

A lean test that strips out the mechanism doesn't test the idea — it buries it.

Difficulty
Easy
Time to result
~weeks to results
Steps
5
Confidence
90%

Gina Gotthilf (early growth leader at Duolingo, now COO of Latitude) deprioritized badges for six months on ROI grounds, then ran the leanest possible version — sign up, get one badge — which produced nothing, so the team moved on for another eight months. The lean test had removed every property that makes badges work: pride in the achievement, a collection to complete, an audience to show. The framework is a checklist that prevents a cheap experiment from generating a false negative on a real idea, plus the dogfooding habit that would have caught it in minutes.

Origin

Gina Gotthilf's account of Duolingo's growth team and the badges experiment — an idea that, once done properly, positively impacted almost every metric at the company, including ones they hadn't expected.

Core principles

  • 01Not every idea can be tested cheaply. Some have a minimum viable mechanism below which the test is meaningless.
  • 02Before running a lean test, list the properties that make the idea work — then check your test preserves them.
  • 03A null result from a mechanism-stripped test is a false negative, and false negatives cost you years.
  • 04ROI-ranked backlogs systematically under-rate ideas whose payoff requires expensive setup.
  • 05Dogfood your experiments, not just your product. Anyone who had used that badge would have known it was lame.
  • 06Some wins are platforms: once people want badges, you can ask them to do anything to earn one.

How to run it

  1. 1

    Write down why the idea should work

    Before designing the test, articulate the mechanism. For badges: people are proud of a real achievement, they want to complete a collection, and they can show it to others. That's the causal engine.

    Pro tip Study the source you copied it from. The growth team played all the popular games — the mechanism was visible in every one of them.

  2. 2

    Check the lean version against the mechanism list

    Design the cheapest test you can, then hold it against the list from step 1 and mark which mechanisms survive. Duolingo's version: badge for signing up (no pride — nobody's proud of signing up), one badge (no collection), invisible to others (no display). Zero of three survived.

    Pro tip If fewer than half your mechanisms survive the lean version, the test cannot produce a valid negative. Either raise the treatment or don't run it.

    Watch out This is the trap: the leanest test is not always a valid test. 'Cheap' and 'informative' are different properties.

  3. 3

    Dogfood the experiment before you launch it

    Go through the experience yourself, as a user, before shipping the test. The badge with the girl and the balloon was self-evidently lame the moment anyone actually saw it in context. Make this a standing practice, not a one-off.

    Pro tip Growth teams built from marketers (not product people) commonly skip dogfooding without knowing the term. Make it explicit in the experiment checklist.

    Watch out Rigorous prioritization and immaculate write-ups do not substitute for one person actually using the thing.

  4. 4

    Interrogate every null result before shelving the idea

    When a test comes back flat, ask which mechanism was missing before you conclude the idea is dead. Duolingo declared badges a failure and moved on for eight more months on the strength of a test that never contained badges' actual value.

    Pro tip Keep a 'false negative' review on the backlog: quarterly, revisit dead ideas and ask whether the test was strong enough to have killed them.

  5. 5

    Weight platform ideas above their direct ROI

    Some ideas are not features but currencies. Once people want badges, you can attach them to any behavior — invite friends, buy things, anything. That optionality never appears in a per-experiment ROI ranking, so it must be added deliberately.

    Pro tip Flag candidate ideas that create a new incentive currency and score them on unlockable follow-on behaviors, not just first-order lift.

    Watch out Ranking purely on estimated users-gained-per-time-invested is exactly what kept badges buried for six months.

In the wild

Duolingo badges

The growth team wanted badges after studying popular games. Because experiments were ranked by ROI (expected users vs time investment) and badges looked like a big time sink, Gotthilf blocked the experiment for about six months. When they finally ran it, they ran the leanest possible version: sign up, receive a single badge — a picture of a girl with a balloon. It produced no results, so they moved on for another eight months. Later they realized nobody is proud of signing up, there was no collection to build, and no one to show it to; none of the things that make badges compelling were present. They also realized the growth team had never been dogfooding.

Done properly, badges positively impacted almost every metric at the company, including some they hadn't expected, and became a platform — once people want badges, you can ask them to do anything to earn one. Dogfooding became a standing practice.

Common mistakes

MVP-ing the mechanism away

The lean version of badges contained no pride, no collection and no audience. The test was cheap and completely uninformative — but the team read its null result as a verdict on the idea.

Not dogfooding your own experiments

The growth team was scrupulous about prioritization and write-ups but never used the experiences it shipped. A single person going through the flow would have seen the badge was lame before it ever ran.

Pure ROI ranking on the experiment backlog

Scoring ideas by expected users-gained over time-invested reliably buries high-setup, high-ceiling ideas. Badges sat for six months and then eight more, on a scoring rule that couldn't see their platform value.

Is it for you?

Best for

Growth teams with an ROI-ranked experiment backlog who need to avoid killing high-ceiling ideas with under-powered lean tests.

Not ideal for

Simple, single-mechanism changes (copy, pricing display, button placement) where the cheap version genuinely is the idea.

From the transcript

we ran this very simple experiment that was like you signed up and then you get a badge

54:00

signing up it's not an exciting moment and you don't even have badges to collect you can't show it to other people like none of…

54:30

we discovered that we hadn't been dog fooding

54:30

why are we not testing our experiments

55:00

From the episode

Failure