Hypothesis-Driven Experimentation: Learning as a Win
Reframe experiments from winners/losers to hypotheses; run more, run shorter, and carry priors forward.
- Difficulty
- Advanced
- Time to result
- ~ongoing to results
- Steps
- 5
- Confidence
- 92%
Cultures that reward data scientists for 'wins' quietly push them toward incremental tests run too long, missing big opportunities in the tails. Johari's fix: make learning an explicit win, budget for the cost of learning, run more and shorter experiments into the tails, and use Bayesian methods so past experiments inform future ones instead of being thrown away.
Origin
Ramesh Johari; draws on a Microsoft paper on 'A/B testing with fat tails' (not his own) and the statistical lineage of Ronald Fisher.
Core principles
- 01Rewarding 'wins' biases teams toward incremental designs run longer than they should be
- 02Experimentation is historically hypothesis-driven, not a win/loss scoreboard
- 03A failed test that was rigorously hypothesis-driven still teaches you a lot
- 04Learning is not free — you must be willing to 'pay to learn' by spending samples on options that may lose
- 05Frequentist tests discard prior evidence; Bayesian A/B testing lets each experiment update a prior for all future ones
How to run it
- 1
Spot the incentive trap
Notice if data scientists are judged by number of wins per quarter; this pushes incremental tests and long runtimes.
Watch out It's easier to get wins when you're being incremental, and you must run those long to prove them — quietly killing risk-taking.
- 2
Redefine every test as a hypothesis, not a win/loss
Require each experiment doc to state what you'll learn about the business — a funnel, demand elasticity, guest/host preferences — regardless of outcome.
Pro tip Leaders should expect data scientists to report what they learned about the business, not just statistically rigorous win/loss results.
- 3
Run more, shorter, into the tails
Increase velocity: try riskier, non-incremental ideas and stop running everything so long, accepting bigger failures for bigger potential wins.
Watch out Don't hit stat-sig and reflexively green-light; one outlier A/B test shouldn't overturn everything you know about the business.
- 4
Budget for the cost of learning
Accept that allocating samples to options that end up losing is the price of knowing which option is better — you'd never give all samples to control if you'd known.
Pro tip A holdout that 'cost' a few million told leadership exactly what the data team was worth — an answer unobtainable any other way.
Watch out If you reward only shipped winners, you implicitly brand all learning-by-losing as wasted time.
- 5
Carry learning forward with Bayesian priors
Encode results from past experiments into a prior so future tests combine prior belief with new data, creating a positive information externality across the org.
Pro tip This lets a 'failed' experiment count as progress because it moved the prior that shapes all future analysis.
In the wild
A real-estate platform's marketing data-science manager secretly held out a group from all the innovations. It cost a couple of million dollars — but revealed the true value of the team's work and gave leadership an answer they could not have gotten otherwise.
→ Made visceral that learning has a price worth paying, and that framing spend on 'losing' arms as waste is a hindsight illusion.
Airbnb's Superhost badge was feared to wreck the finely tuned ranking; the A/B test showed essentially no short-term business impact, only that hosts felt more satisfied. Rather than 'losing,' it taught how attention reallocates and hinted the real value was long-term host retention.
→ A flat headline metric still carried a hypothesis-level lesson; the feature was retained and judged valuable beyond measurable short-run lift.
Common mistakes
Judging data scientists by win count
It selects for incremental tests and over-long runtimes, starving the org of big-swing learning in the tails.
Throwing away the past every experiment
Frequentist analysis ignores the thousand prior tests on the same button; Bayesian priors recover that information.
Treating stat-sig as an automatic green light
Data science is accumulation of evidence; a lone outlier result shouldn't overturn everything you believe about the business.
Is it for you?
Best for
Heads of data science and growth leaders designing experimentation culture and incentives in an experiment-heavy org
Not ideal for
Teams without enough traffic or data maturity to run and analyze many experiments
From the transcript
“it's easier to get wins when you're being incremental”
“so learning is a win”
“that that idea that you have to pay to learn”
“that falls broadly under the category of what's called basian AB testing”
“I like to push people to be what's called Quantified rather than data driven”
From the episode
Marketplace lessons from Uber, Airbnb, Bumble, and more
Ramesh Johari (Stanford professor, startup advisor)