LLenny's Podcast
← All frameworks
MindsetRonny Kohavi (Airbnb, Microsoft, Amazon)

Twyman's Law: Investigate the Too-Good-to-Be-True Result

Any surprising or extreme experiment result is probably wrong until you've hunted for the flaw

Difficulty
Easy
Time to result
~days to results
Steps
4
Confidence
95%

Twyman's Law states that any figure that looks interesting or different is usually wrong. Applied to experimentation, a result far outside your normal movement range should trigger investigation, not celebration, because most enormous results turn out to be measurement flaws.

Origin

Twyman's Law originates with Tony Twyman, a UK media/radio audience-measurement expert; Kohavi adopted and popularized it for online experimentation.

Core principles

  • 01A surprising result is one where the pre-experiment estimate and the actual result differ by a large absolute amount
  • 02Surprising losers teach as much as surprising winners — investigate both
  • 03When a result is far outside normal movement (e.g. 10% when you usually see under 1%), suspect a bug first
  • 04Roughly nine out of ten dramatic results, when investigated, reveal a flaw
  • 05Genuine breakthroughs survive replication and double/triple-checking

How to run it

  1. 1

    Set your normal-movement baseline

    Know the typical magnitude of change your experiments produce so you can recognize an outlier.

  2. 2

    Flag results that break the baseline

    When a result is dramatically larger (or more negative) than usual, 'hold the celebratory dinner' and treat it as suspect.

    Pro tip Your internal disbelief meter should rise in proportion to how extreme the number is.

    Watch out People want to see success — a natural bias pushes teams to accept good-looking results uncritically.

  3. 3

    Hunt for the flaw before acting

    Look for logging errors, pipeline issues, bots, or sample-ratio mismatch; only trust the result once you can't find a problem.

    Watch out Historically an alarm firing on a big revenue swing meant a real bug (double-logged revenue), so a huge move is a prior toward error.

  4. 4

    Replicate genuine surprises

    If no flaw is found, re-run the experiment to confirm before broadcasting the win.

    Pro tip Only broadcast a surprising win after independent replication has confirmed it.

In the wild

The Bing ad-title revenue alarm

Moving the second line of an ad up to enlarge the title line increased revenue about 12%. A revenue-anomaly alarm fired, which had previously always meant a bug, so the team's first reaction was 'this is too good to be true, let's find the bug' and they replicated it several times.

No bug existed — it was the largest revenue win in Bing's history (worth ~$100M), validated precisely because it survived repeated checking.

Windows indexer battery drain

A Windows search-indexer change showed higher offline relevance and looked like a clear win, but the live experiment surfaced a surprising negative from left field.

The change killed laptop battery life by consuming far more CPU; investigating the surprising loser produced a new factor to design around.

Common mistakes

Celebrating a big result before investigating

Taking the team to dinner over a 10x-normal result usually precedes discovering the number was a bug, wasting credibility and momentum.

Only scrutinizing wins, not surprising losses

Ignoring unexpectedly negative results forfeits the deepest learning, which often comes from understanding why a 'no-brainer' idea tanked.

Is it for you?

Best for

Analysts and PMs interpreting experiment scorecards who need a discipline against confirmation bias

Not ideal for

Low-stakes contexts where results are small and within normal variance and no dramatic outliers appear

From the transcript

the general statement is if any figure that looks interesting or different is usually wrong

1:00:30

if the result looks too good to be true if you suddenly moved your you know your normal movement of an experiment is under one…

1:01:00

nine out of ten when we call out time is law it is the case that we find some flaw in the experiment

1:01:30

From the episode

The ultimate guide to A/B testing

Ronny Kohavi (Airbnb, Microsoft, Amazon)