LLenny's Podcast
← All frameworks
InnovationRonny Kohavi (Airbnb, Microsoft, Amazon)

Incremental OFAT Over Big-Bang Redesign

Decompose big redesigns into one-factor-at-a-time tested steps because ~80% of ideas fail

Difficulty
Moderate
Time to result
~months to results
Steps
4
Confidence
94%

Because most ideas fail (66-92% depending on domain maturity), bundling 17 changes into a single redesign is statistically likely to be net-negative. The remedy is to move incrementally — one factor at a time (OFAT) — testing and adjusting on the way, so you keep the few good changes and drop the losers instead of shipping them all bundled.

Origin

Kohavi's articulation of one-factor-at-a-time (OFAT) experimentation applied to product redesigns, drawn from his experience at Microsoft, Amazon, and Airbnb.

Core principles

  • 01Two-thirds to over 90% of ideas fail to improve the target metric, depending on how optimized the domain already is
  • 02Bundling many changes makes a negative outcome more likely and hides which factors helped or hurt
  • 03Move incrementally in a believed-good direction, testing and adjusting as you go
  • 04Sunk cost ('we spent six months, we can't unlaunch') is the trap that ships known-bad redesigns
  • 05Big redesigns aren't forbidden — but allocate them as high-risk bets and be ready to fail ~80% of the time

How to run it

  1. 1

    Try to decompose the redesign

    Break the intended redesign into individual factors or a small set of factors that can each be tested.

    Watch out If you can't decompose it at all, do a small set of factors at a time rather than all 17 at once.

  2. 2

    Test one factor at a time and learn

    Ship each factor as its own experiment, learn from the result, and adjust your direction.

    Pro tip Of 17 changes you may find only ~4 good ideas — those are the ones worth launching.

  3. 3

    Keep winners, drop losers, and don't ship flat

    Launch only the factors that move the metric positively; refuse to ship flat or negative changes just because effort was spent.

    Pro tip A flat result is a no-ship: shipping it adds code and maintenance cost for zero value.

    Watch out Exception: legal requirements may force shipping flat/negative — then test variants and ship the one that hurts least.

  4. 4

    Budget a portion for high-risk redesigns

    Allocate a deliberate share of effort (roughly a 70/20/10 known/bet/infrastructure split) to big redesigns, entering them knowing ~80% will fail.

    Pro tip Frame the big bet honestly: 'most likely you'll fail, but if it wins it's a breakthrough.'

    Watch out Don't let passion convince the team the redesign is the exception — organizations that run experiments get humbled early.

In the wild

Airbnb search relevance, 250 experiments

Over Kohavi's tenure the search-relevance team ran ~250 experiments where 92% failed to move the target metric, yet the successful 8% compounded into a ~6% revenue improvement — no single idea, but many small tested gains.

Incremental tested changes produced a large aggregate win that a single bundled redesign would have risked as one net-negative launch.

Bing social integration big bet

Bing spent roughly a hundred person-years integrating the Twitter firehose and Facebook feed as a big strategic bet; across hundreds of experiments the results were negative to flat.

The feature was aborted after ~18 months — a fair bet that the data eventually killed, illustrating why big bets need a fail-ready posture.

Common mistakes

Bundling many changes into one redesign

With most individual ideas failing, combining 17 untested changes makes the whole launch more likely to be negative and obscures which factors caused the harm.

Shipping a flat or negative result due to sunk cost

Launching because 'the team will be demotivated' or 'we already spent the time' adds maintenance burden and hurts users for no measurable gain.

Is it for you?

Best for

Product teams tempted by a full redesign or onboarding-flow rebuild who want to de-risk it

Not ideal for

Truly indivisible changes, or organizations without enough traffic to test factors independently

From the transcript

do them in smaller increments learn from it's called o-fat one factor at a time do one factor learn from it and adjust of the…

38:00

if you believe in that statistics that I published then doing 17 changes together is more likely to be negative

37:30

you don't ship on flat unless it's a sort of a legal requirement

44:30

From the episode

The ultimate guide to A/B testing

Ronny Kohavi (Airbnb, Microsoft, Amazon)