LLenny's Podcast
← All episodes
Ronny Kohavi (Airbnb, Microsoft, Amazon)27 July 2023

The ultimate guide to A/B testing

7Frameworks
15Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster1:02:00

What a P-Value Actually Means (and Why 0.05 Isn't 5%)

Kohavi corrects the most common p-value error: people read 1 minus the p-value as the probability the treatment beats control, which is wrong. The p-value assumes the null hypothesis is true; getting the number you actually want requires Bayes' rule and a prior. Given Airbnb search's 8% success rate, a result at p < 0.05 carries a 26% chance of being a false positive — not 5% — which is why he pushed teams to p < 0.01 plus replication.

  • People wrongly read 1 minus p-value as the probability treatment beats control
  • A p-value assumes the null hypothesis is true; it's a conditional probability
  • Recovering the desired probability needs Bayes' rule plus a prior success rate
  • At Airbnb search's 8% success rate, a p < 0.05 result is 26% likely to be a false positive
  • Fix: require p < 0.01 and replicate, combining experiments via Fisher's/Stouffer's method

many people assign one minus p-value as the probability that your treatment is better than control

Ronny Kohavi · 1:02:30

if you get a statistically significant result with the p-value less than 0.05 there's a 26 percent chance that this is a false positive result…

Ronny Kohavi · 1:04:30
#ab-testing#p-value#statistics#false-positive

Hot Take· 3

Hot Take20:30

You Can't Experiment Too Much — But Balance Your Portfolio

Kohavi is a strong advocate of 'test everything' — any code change or feature should sit inside an experiment, because even small bug fixes can have surprising impact. He argues you can't over-experiment, but you can over-focus on incremental changes. Like a stock portfolio, you need a mix of safe incremental bets and high-risk, high-reward swings that will usually fail.

  • Every code change or feature should be inside an experiment
  • Even small bug fixes can have surprising, unexpected impact
  • It is not possible to experiment too much
  • You should allocate some effort to high-risk, high-reward bets
  • Be ready to fail ~80% of the time on the big swings

and so I don't think it's possible to experiment too much

Ronny Kohavi · 21:00

it's like in stock you need a portfolio you need some experiments that are incremental that move you in the direction

Ronny Kohavi · 21:30
#ab-testing#experimentation-culture#risk#portfolio
Hot Take36:00

Big Redesigns Almost Always Fail

Both agree that full product redesigns rarely test positive, and teams then have to claw back what they broke. Kohavi's advice: if you must redesign, do it in steps and test along the way rather than shipping 17 changes at once, since combining many likely-to-fail changes makes the whole more likely to be negative. He calls the incremental approach OFAT — one factor at a time.

  • Full redesigns almost never produce a positive result
  • Doing 17 changes together is more likely to be negative
  • Redesign in steps and adjust as you learn
  • OFAT — one factor at a time — isolates what works
  • The sunk-cost fallacy pressures teams to ship redesigns even when they hurt users

so the right way to do this is to say yes we want to do a redesign but let's do it in steps and test…

Ronny Kohavi · 36:30

do them in smaller increments learn from it's called o-fat one factor at a time do one factor learn from it and adjust

Ronny Kohavi · 38:00
#ab-testing#redesign#incremental#ofat
Hot Take44:00

Never Ship a Flat Result

Kohavi argues that if you're not experimenting, roughly 70% of what you ship is hurting the business — some flat, some negative. A flat result should be a no-ship, because new code adds maintenance overhead with no offsetting value. The only exception is a legal requirement, and even then you should test several options and ship the one that hurts least.

  • Without experiments, ~70% of shipped work is flat to negative
  • A statistically-flat result is a no-ship: new code carries maintenance cost
  • 'The team will be demotivated if we don't ship' is the wrong reason to ship
  • Legal requirements are the exception where you may ship on flat or negative
  • Even for legal, test options and ship the least harmful

if you're not writing an experiment 70 of stuff you're shipping is hurting your business

Ronny Kohavi · 44:00

you don't ship on flat unless it's a sort of a legal requirement

Ronny Kohavi · 44:30
#ab-testing#shipping#maintenance#decision-making

Explainer· 4

Explainer10:30

Big Wins Are Rare; Progress Comes Inch by Inch

Kohavi warns that the one-hour-of-work-for-a-massive-win 'gold nugget' is very rare. Most gains accumulate through many tiny improvements: Bing's relevance team targets a 2% annual lift built from many 0.1–0.2 gains, and at Airbnb search relevance ~250 experiments added up to a 6% revenue improvement — even though 92% of the ideas failed to move the target metric.

  • Bing relevance aims for ~2% improvement per year, assembled from tiny gains
  • Airbnb search relevance ran ~250 experiments in his tenure
  • Those small wins summed to a 6% revenue improvement
  • 92% of those experiments failed to move the target metric
  • Only ~8% of ideas actually moved the key metrics

but yes I think most results are inch by inch you improve small amounts lots of them

Ronny Kohavi · 11:30

of these experiments 92 percent failed to improve the metric that we were trying to move so only eight percent of our ideas actually were…

Ronny Kohavi · 12:30
#ab-testing#incremental-gains#airbnb#bing
Explainer13:00

How Often Experiments Actually Fail: 66% to 92%

Kohavi shares real failure rates across his career: roughly two-thirds of ideas failed at Microsoft, ~85% at the more optimized Bing, and 92% at Airbnb — the highest he's observed. Other companies like Google Ads and Booking publish 80–90% failure rates. He notes these are rates of experiments, not ideas, since ~10% of experiments are aborted on day one for implementation bugs.

  • Microsoft: about two-thirds of ideas fail
  • Bing: ~85%, because a well-optimized domain is harder to improve
  • Airbnb: 92%, the highest he has observed
  • Google Ads, Booking and others publish 80–90% failure rates
  • ~10% of experiments are aborted on the first day for implementation issues, not bad ideas

so overall at Microsoft about 66 two-thirds of ideas fail right

Ronny Kohavi · 13:30

and then at Airbnb uh this 92 number is you know the highest failure rate that I've observed

Ronny Kohavi · 14:00
#ab-testing#failure-rate#benchmarks
Explainer55:30

Sample Ratio Mismatch: The Most Common Way Experiments Go Wrong

Kohavi's most common validity red flag is a sample ratio mismatch: if a 50/50 design comes back off (say 50.2/49.8 on a million users), a formula can show that split should occur only once in half a million experiments — so something is wrong, often bots or a data pipeline. Microsoft found ~8% of experiments suffered from it, and Kohavi tells a funny story of red-lining every scorecard number so screenshots couldn't hide the flag.

  • A 50/50 design that returns off-ratio signals something is broken
  • A formula computes the probability the split happened by chance
  • Bots and data-pipeline issues are common causes
  • ~8% of Microsoft experiments suffered a sample ratio mismatch
  • Product managers ignored banners, so the team red-lined every scorecard number for screenshots

but I'll say maybe one of the things that is the most common occurs by far which is a sample ratio mismatch

Ronny Kohavi · 55:30

every number in the scorecard was highlighted with a red line so that if you took a screenshot other people could tell you that a…

Ronny Kohavi · 1:00:00
#ab-testing#validity#sample-ratio-mismatch#bots
Explainer1:00:30

Twyman's Law: If It Looks Too Good, It's Probably Wrong

Kohavi explains Twyman's Law — any figure that looks interesting or different is usually wrong. When a metric that normally moves under 1% suddenly jumps 10%, hold the celebration and investigate, because nine out of ten times a flaw turns up. Genuine outliers exist (like the ad-title win) but survive only after repeated replication and double-checking.

  • Twyman's Law: any figure that looks interesting or different is usually wrong
  • A sudden 10% move where movement is normally under 1% is a warning, not a win
  • Nine out of ten times investigating reveals a flaw
  • Real outliers exist but must be replicated and double-checked
  • Fight the natural bias to want to see success

the general statement is if any figure that looks interesting or different is usually wrong

Ronny Kohavi · 1:00:30

nine out of ten when we call out time is law it is the case that we find some flaw in the experiment

Ronny Kohavi · 1:01:30
#ab-testing#twymans-law#skepticism#validity

Story· 5

Story05:00

The Two-Line Swap That Made Bing $100M

Kohavi's favorite public example: an engineer, tired of an idea languishing in the backlog, spent a couple of days moving an ad's second line up to the larger title line. The experiment tripped a revenue alarm because it lifted revenue ~12% without hurting user guardrail metrics — the biggest revenue win in Bing's history from a trivial change. It illustrates how badly teams predict experiment outcomes and how ROI-driven quick tests uncover them.

  • The idea sat in the backlog for months, rated below many others
  • Implementation took only a couple of days; an alarm fired for a suspected revenue bug
  • Revenue rose ~12% with no bug found after repeated replication
  • Worth ~$100M at the time, and it did not hurt the user experience guardrails
  • Teams are routinely humbled at predicting which experiments win

this thing was worth a hundred million dollars at the time when Bing was a lot smaller and the key thing is it didn't hurt…

Ronny Kohavi · 07:00

but to me this was an example of a tiny change that was the best Revenue generating idea in Bing's history

Ronny Kohavi · 08:30
#ab-testing#bing#surprising-results#revenue
Story08:30

Opening Links in a New Tab Keeps Winning

Both host and guest describe the same surprising winner appearing across companies: opening a clicked search result in a new tab. Kohavi first ran it around 2008 for Hotmail and MSN despite heavy designer pushback, and it was highly beneficial. At Airbnb the win had been discovered, then semi-forgotten, until it was reintroduced — a lesson in institutional memory.

  • Opening results in a new tab was one of the biggest search wins at Airbnb
  • Kohavi ran the same test ~2008 for Hotmail then MSN, with strong results
  • Designers pushed back because users didn't ask for it
  • Airbnb had implemented it, forgot the learning, then reintroduced it
  • Winners must be documented so the organization remembers them

the search team just ran a small experiment of what if we were to open a new tab every time someone clicked on a search…

Lenny Rachitsky · 09:00

which is one of the things you learned about institutional memories when you have winners make sure to address them and remember them

Ronny Kohavi · 10:00
#ab-testing#institutional-memory#search#airbnb
Story22:30

Bing's Big Social Bet That Quietly Died

A cautionary tale of a large, high-conviction bet: Bing tried to reshape search by integrating social, hooking into the Twitter firehose and Facebook. After roughly a hundred person-years and hundreds of experiments, all results were negative to flat, and the feature was aborted after about a year and a half. Kohavi notes similar social attempts failed at Netflix and Airbnb.

  • Bing hooked into the Twitter firehose and Facebook to integrate social
  • Roughly a hundred person-years were spent on the idea
  • Hundreds of experiments were all negative to flat
  • The feature existed ~18 months before being aborted
  • Netflix and Airbnb had similar failed social attempts

we hooked into the Tweeter fire hose feed and we hooked it to Facebook and we spent a hundred percent years on this idea and…

Ronny Kohavi · 23:00

it's a million dollar question you know you continue and then maybe the breakthrough will come next month

Ronny Kohavi · 23:30
#ab-testing#big-bets#social#bing
Story33:00

How Amazon's Email Team Gamed Its Own Metric

At Amazon, the recommendation-email team was credited for any purchase that followed an email — a metric with no countervailing cost, so they simply sent more emails and 'made more money.' Kohavi's team modeled the long-term cost of an unsubscribe (a few dollars of lost lifetime value), and found more than half the campaigns were actually negative. The insight led to per-campaign unsubscribe defaults.

  • The email team got credit for any purchase following an email
  • With no countervailing metric, more emails meant more claimed revenue — leading to spam
  • A data science study priced the lost lifetime value of an unsubscribe
  • Once modeled, more than half the campaigns were net negative
  • The fix: default to unsubscribing from that specific campaign, not all email

the more emails you send the more money you're going to credit the theme and so that led to spam literally

Ronny Kohavi · 34:00

when we started to incorporate those formula more than half the campaigns that were being sent were negative

Ronny Kohavi · 35:00
#ab-testing#email#lifetime-value#amazon
Story53:00

How Optimizely Lost Trust With Real-Time P-Values

Kohavi uses early Optimizely as a cautionary tale about trust. Their real-time approach let users stop an experiment the moment the p-value hit significance — a practice that inflates the false positive rate from a nominal 5% to around 30%. Customers saw 'wins' that never showed up in real revenue, a famous 'Optimizely almost got me fired' post surfaced, and the company eventually fixed the statistics.

  • Early Optimizely computed p-values in real time and let users stop at significance
  • Continuous monitoring inflates the type-one (false positive) error rate
  • A nominal 5% error rate could balloon to ~30%
  • Customers' reported wins didn't materialize in actual revenue
  • After public criticism, Optimizely corrected its methodology

optimizely in its early days were very statistically naive they sort of said hey we're real time we can compute your P values in real…

Ronny Kohavi · 53:30

using real-time sort of p-value monitoring to optimize the offered you would probably have a 30 percent error rate

Ronny Kohavi · 54:00
#ab-testing#trust#statistics#optimizely

Tool· 1

Tool15:30

goodui.org: A Library of Patterns That Often Win

Asked for a list of changes that tend to work, Kohavi points to two resources: Microsoft's 'Rules of Thumb' paper extracting patterns from thousands of experiments, and goodui.org by Jakub Linowski. goodui.org crowdsources experiment results into ~140 patterns, each showing how often it helped and by how much — even the 'open in a new window' win appears there.

  • Microsoft's 'Rules of Thumb' paper extracted patterns from thousands of experiments
  • goodui.org collects submitted results into named patterns
  • ~140 patterns, each with a win rate and effect size
  • The 'open a new window' pattern appears in it
  • Kohavi suggests these patterns could feed an AI-generated roadmap

but there's another more more accurate I would say uh resource that's useful that I recommend to people and it's a site called goodui.org

Ronny Kohavi · 16:00

he puts them into patterns there's probably like 140 patterns I think at this point

Ronny Kohavi · 16:30
#ab-testing#tools#patterns#roadmap

Takeaway· 1

Takeaway26:30

The 200,000-User Rule for When to Start A/B Testing

Kohavi's practical heuristic for when experimentation starts paying off: below tens of thousands of users the statistics don't work for most metrics. To detect the 5–10% improvements startups should target on a typical retail conversion rate, you need around 200,000 users. Below that, start building the culture and platform so value compounds as you scale.

  • Below tens of thousands of users the math doesn't work for most metrics
  • Startups should chase 5–10% effects, not 1% ones
  • Detecting a ~5% retail conversion lift needs roughly 200,000 users
  • At 10s of thousands you can only detect large effects
  • Below the threshold, build culture, platform and integration for later

unless you have at least tens of thousands of users the math the statistics just don't work out for most of the metrics that you're…

Ronny Kohavi · 27:00

so you ask for rule of thumb 200 000 users you're magical below that start building the culture start building the platform

Ronny Kohavi · 27:30
#ab-testing#startups#statistics#rule-of-thumb