Conviction-Building for Low-Volume Experiments
When you can't run a clean A/B test, stack alternative signals to raise conviction instead of faking precision
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 4
- Confidence
- 90%
A decision framework for teams that can't reach statistical significance because transaction volume is too low. You first run a power analysis and honestly accept the runtime; if a clean A/B test isn't feasible, you reframe the goal as increasing conviction and reach for a menu of weaker-but-usable techniques; and if none work, you trust intuition and ship — without pretending you have precision you don't.
Origin
Brian Tolkin's approach at Opendoor, where the business does far fewer, far larger transactions than Uber's millions per second.
Core principles
- 01Experimentation is fundamentally about increasing your conviction in the problem or solution
- 02Acknowledge the problem: run a power analysis before forcing yourself into an A/B test
- 03Some experiments are important enough that a six-month runtime is an acceptable, deliberate choice
- 04The only real mistake is thinking you'll get an answer in a month when you won't, then pretending you did
- 05Don't spend time trying to get false precision
How to run it
- 1
Run a power analysis first
Before committing to an A/B test, compute the minimum detectable effect, the sample size, and the runtime. Use a calculator that lets you plug in traffic and acceptable runtime and returns the minimum detectable impact, then gut-check it against intuition.
Pro tip If a calculator gives you a minimum detectable effect larger than any plausible real effect, the test is dead on arrival — know that up front.
Watch out Top-of-funnel tests are easier than down-funnel; feature/tech tests are easier than operational-process tests.
- 2
Decide honestly whether the runtime is acceptable
If the important-enough experiment needs six months, deliberately choose to 'set it and forget it' and be grateful you started early (e.g. start in June to be smarter for next-year planning).
Watch out Never assume a one-month answer on an experiment that mathematically needs longer, then wake up a month later calling it 'insignificant.'
- 3
If no clean test is possible, stack alternative conviction-builders
Reframe as 'how else can I increase conviction?' Options: talk to more customers (best/most obvious), use observational data, compare sister/twin cities, segment by geo, reduce power (run at 80% confidence instead of 95% as a worthy tradeoff), or run a long-term holdout.
Pro tip Running at 80% confidence means being wrong one more time out of ten — often an acceptable tradeoff for low-volume flows.
Watch out These techniques aren't as rigorous as a canonical A/B test — treat them as conviction-raisers, not proof.
- 4
If nothing works, trust intuition and ship — with a feedback loop
When there's no significance and no other technique, use taste and judgment and ship. But wire up a reasonable feedback loop (support ticket volume, future adoption) to learn whether you were right.
Pro tip Before shipping on intuition, rate your conviction low/medium/high; for medium-or-low conviction on consequential decisions, talk to more customers or gut-check with another person to reach high.
In the wild
Because Opendoor runs relatively few, large transactions, some experiments can't hit significance quickly. Rather than pretend, the team will accept a six-month runtime on an important experiment, start it in June, and 'set it and forget it' to be smarter for the next planning cycle.
→ Avoids the trap of declaring premature insignificance and preserves a real answer for future planning.
Common mistakes
Forcing yourself into an A/B test without a power analysis
You launch a test you can never resolve, then a month later call the result insignificant when you could have known that in advance.
Chasing false precision on inherently low-signal flows
When no rigorous method exists, over-investing in statistical theater wastes time; sometimes intuition plus a feedback loop is the honest answer.
Is it for you?
Best for
Product and growth teams at low-volume, high-value businesses (real estate, enterprise, B2B) where clean A/B tests are often infeasible
Not ideal for
High-volume consumer products where standard A/B testing already yields fast significance
From the transcript
“don't just for for yourself into AB testing without running the power analysis”
“there are certain experiments that are important enough ... that you may say six-month runtime is an acceptable outcome”
“the only mistake here is like thinking you'll get an answer in a month when you won't and then pretending you do”
“we're going to run at 80% confidence for all of our experiments instead of the traditional 95% because that's a worthy tradeoff and if we're…”
“you shouldn't spend time trying to get false Precision”
From the episode
Lessons from scaling Uber and Opendoor
Brian Tolkin (Head of Product at Opendoor, ex-Uber)