LLenny's Podcast
← All frameworks
StrategyRonny Kohavi (Airbnb, Microsoft, Amazon)

The 200K-User Experimentation Readiness Rule

Below ~200K users detect only big effects; use the wait to build platform and culture

Difficulty
Easy
Time to result
~months to results
Steps
4
Confidence
92%

A practical-defaults heuristic for when a company can meaningfully A/B test. Below tens of thousands of users the statistics don't work for most metrics; around 200,000 users you can reliably detect the 5-10% effects startups should care about. Below that threshold, invest in culture and platform so value compounds as you scale.

Origin

From Ronny Kohavi's 'Practical Defaults' talk, giving concrete sample-size guidance for startups considering experimentation.

Core principles

  • 01Below tens of thousands of users the statistics don't work out for most metrics of interest
  • 02Startups should chase 5-10% effects, not 1% effects, because only large effects are detectable at small scale
  • 03A retail conversion example needs ~200,000 users to detect a ~5% beneficial change
  • 04You need enough units (usually users) for randomization statistics to hold — some domains (M&A) can never be A/B tested
  • 05The pre-threshold period is for building the culture, platform, and integrations, not idle waiting

How to run it

  1. 1

    Check you have enough units

    Confirm your experimental unit (usually users) exists in sufficient volume; some decisions (mergers, one-off acquisitions) inherently can't be A/B tested.

    Watch out Too few units means the statistics simply won't resolve real effects.

  2. 2

    Below tens of thousands, focus only on large effects

    If you're in the tens of thousands of users, accept you can only detect large effects and aim experiments at 5-10% swings.

    Pro tip Don't waste small-scale traffic hunting for 1% wins you can't statistically see.

  3. 3

    At ~200K users, start testing broadly

    Once you reach roughly 200,000 users, the 'magic starts' — you can test much more and guard against degradations.

  4. 4

    Use the pre-threshold time to build

    Before you hit the threshold, build the experimentation culture, platform, and integrations so value appears the moment you scale.

    Pro tip Consult anyone in the org with prior experimentation experience while you build.

In the wild

Retail conversion sample-size math

For a retail site trying to detect changes of at least ~5% in conversion rate, Kohavi's practical number was roughly 200,000 users needed for the statistics to work.

Gives founders a concrete go/no-go threshold instead of testing prematurely and drawing noise-based conclusions.

Common mistakes

Running A/B tests before you have the traffic

With too few users the statistics can't resolve real effects, so teams chase noise and draw false conclusions from underpowered tests.

Chasing 1% effects at small scale

Small-scale traffic can only reveal large effects, so hunting tiny improvements wastes the limited statistical power available.

Is it for you?

Best for

Startup founders and growth leads deciding whether it's time to stand up experimentation

Not ideal for

Decisions with too few units to randomize (M&A, one-off strategic bets) which can never be A/B tested

From the transcript

unless you have at least tens of thousands of users the math the statistics just don't work out for most of the metrics that you're…

26:30

they shouldn't focus on the one percent they should focus on the five and ten percent then you need something like 200 000 users

27:00

so you ask for rule of thumb 200 000 users you're magical below that start building the culture start building the platform

27:30

From the episode

The ultimate guide to A/B testing

Ronny Kohavi (Airbnb, Microsoft, Amazon)