The 200K-User Experimentation Readiness Rule
Below ~200K users detect only big effects; use the wait to build platform and culture
- Difficulty
- Easy
- Time to result
- ~months to results
- Steps
- 4
- Confidence
- 92%
A practical-defaults heuristic for when a company can meaningfully A/B test. Below tens of thousands of users the statistics don't work for most metrics; around 200,000 users you can reliably detect the 5-10% effects startups should care about. Below that threshold, invest in culture and platform so value compounds as you scale.
Origin
From Ronny Kohavi's 'Practical Defaults' talk, giving concrete sample-size guidance for startups considering experimentation.
Core principles
- 01Below tens of thousands of users the statistics don't work out for most metrics of interest
- 02Startups should chase 5-10% effects, not 1% effects, because only large effects are detectable at small scale
- 03A retail conversion example needs ~200,000 users to detect a ~5% beneficial change
- 04You need enough units (usually users) for randomization statistics to hold — some domains (M&A) can never be A/B tested
- 05The pre-threshold period is for building the culture, platform, and integrations, not idle waiting
How to run it
- 1
Check you have enough units
Confirm your experimental unit (usually users) exists in sufficient volume; some decisions (mergers, one-off acquisitions) inherently can't be A/B tested.
Watch out Too few units means the statistics simply won't resolve real effects.
- 2
Below tens of thousands, focus only on large effects
If you're in the tens of thousands of users, accept you can only detect large effects and aim experiments at 5-10% swings.
Pro tip Don't waste small-scale traffic hunting for 1% wins you can't statistically see.
- 3
At ~200K users, start testing broadly
Once you reach roughly 200,000 users, the 'magic starts' — you can test much more and guard against degradations.
- 4
Use the pre-threshold time to build
Before you hit the threshold, build the experimentation culture, platform, and integrations so value appears the moment you scale.
Pro tip Consult anyone in the org with prior experimentation experience while you build.
In the wild
For a retail site trying to detect changes of at least ~5% in conversion rate, Kohavi's practical number was roughly 200,000 users needed for the statistics to work.
→ Gives founders a concrete go/no-go threshold instead of testing prematurely and drawing noise-based conclusions.
Common mistakes
Running A/B tests before you have the traffic
With too few users the statistics can't resolve real effects, so teams chase noise and draw false conclusions from underpowered tests.
Chasing 1% effects at small scale
Small-scale traffic can only reveal large effects, so hunting tiny improvements wastes the limited statistical power available.
Is it for you?
Best for
Startup founders and growth leads deciding whether it's time to stand up experimentation
Not ideal for
Decisions with too few units to randomize (M&A, one-off strategic bets) which can never be A/B tested
From the transcript
“unless you have at least tens of thousands of users the math the statistics just don't work out for most of the metrics that you're…”
“they shouldn't focus on the one percent they should focus on the five and ten percent then you need something like 200 000 users”
“so you ask for rule of thumb 200 000 users you're magical below that start building the culture start building the platform”
From the episode
The ultimate guide to A/B testing
Ronny Kohavi (Airbnb, Microsoft, Amazon)