LLenny's Podcast
← All frameworks
StrategyKristen Berman (Irrational Labs)

Evidence-First Intervention Design

Literature review, then hypothesis, then 30 variants, then relative testing — never show users one design.

Difficulty
Advanced
Time to result
~weeks to results
Steps
5
Confidence
90%

Irrational Labs' research process for arriving at an intervention worth shipping. Instead of interviewing five users and guessing, you spend a day in the published literature, form a hypothesis, generate ~30 implementations of it, test five of them against each other on a research panel, then de-risk the in-product experiment down to the two conditions plus control you actually get to run.

Origin

Kristen Berman / Irrational Labs, described in detail as the process behind the TikTok misinformation work. It reflects her origin insight at Intuit: as a PM she was interviewing 5-10 customers to infer human behavior, while an entire academic field had already studied it.

Core principles

  • 01You are not the first person to think about this problem — the literature already contains the answer or the constraint.
  • 02A hypothesis is just a starting point, not a design.
  • 03Never run a study showing people one thing: they may like it, they may hate it, and they may like or hate every design you have — you'd have no idea.
  • 04Panel research doesn't prove market impact; it drives your intuition up before you spend your one shot in-product.
  • 05Behavior is contextual — you cannot drag and drop a winning tactic from another company's context.

How to run it

  1. 1

    Spend a day on Google Scholar before talking to five users

    Search the published literature for what's already been tried on your problem. Aim for a meta-analysis that summarizes the field.

    Pro tip The hot tip is keywords: learn the field's buzzwords (e.g. 'chronic care' in healthcare) or your search returns nothing and you conclude wrongly that nothing exists.

  2. 2

    Extract the mechanism and the timing constraints

    Pull out not just what works but when. For TikTok this yielded two findings: reminding people of their own value of accuracy reduces sharing — and only if delivered at the point of sharing, not before or after.

  3. 3

    Form a hypothesis, then generate ~30 implementations of it

    One hypothesis supports many designs. Irrational Labs produced roughly 30 different ways to implement the TikTok hypothesis before narrowing.

    Watch out Stopping at one implementation of a good hypothesis is how good hypotheses get falsely disproven.

  4. 4

    Test variants against each other on a research panel

    Put five versions in front of 1,000+ users via a panel platform like Prolific. Measure conditions relatively — which condition beats which — never 'do people like this?'

    Pro tip Always present multiple options in a UX study. A single-option study cannot distinguish 'they like it' from 'they'd like anything'.

  5. 5

    De-risk the in-product experiment down to your winners

    You typically get very few conditions in the real product and may not get a second shot. Irrational Labs ended up shipping two conditions plus a control at TikTok — chosen because the panel research had already raised confidence in them.

In the wild

The TikTok share-reduction study

Literature review surfaced the accuracy-prompt and hot-state research. From that hypothesis the team generated ~30 implementations, ran five pop-up versions in front of over a thousand Prolific users measuring conditions against each other, and carried the two strongest plus a control into the live product.

24% reduction in shares of flagged content, a global rollout, and one of the first published studies demonstrating the effect.

The fintech budgeting experiment

Rather than shipping the most-requested feature outright, the team ran a three-condition experiment across 10,000 users — a control plus two distinct budgeting implementations.

A trustworthy null result. Because two implementations were tested, the team could conclude the concept failed, not just one design of it.

Common mistakes

Talking to five users as your research

Berman's founding insight: as a PM she was inferring human behavior from 5-10 customer conversations while a whole discipline had already run the studies. Customer discovery tells you about your users; the literature tells you about humans.

Showing users a single design

A one-condition study confounds liking this design with liking any design. Relative comparison is the only thing a small panel can honestly tell you.

Copy-pasting another company's winning tactic

Behavior is contextual. Berman's team is religious about testing precisely because interventions don't drag-and-drop between contexts — a 133% lift somewhere else is a hypothesis for you, not a result.

Is it for you?

Best for

Teams with one expensive shot at an in-product experiment who need to maximize the odds the shipped variant is the right one.

Not ideal for

Cheap, reversible UI tweaks where running the A/B test directly costs less than the research process.

From the transcript

yeah it's actually a pretty involved process so first we did a literature review which is basically say like look we're not the first people…

32:00

we came up with a hypothesis and we had probably 30 different ways to implement this because the hypothesis is just a starting point and…

32:30

we measured a condition against another condition so relatively what condition is more likely to work than the other condition

33:00

we never do a study a ux study where we're just showing people one thing because they

34:00

we would basically say look go spend a day on the internet Googling to see what else has been done

33:30

From the episode

Using behavioral science to improve your product

Kristen Berman (Irrational Labs)