LLenny's Podcast
← All frameworks
StrategyAmol Avasare

The Two-Bucket Controversial-Test Triage

Sort every risky experiment into a red-line 'never run' bucket or a 'run it if the return justifies the cringe' bucket

Difficulty
Moderate
Time to result
~ongoing to results
Steps
4
Confidence
85%

Growth teams are pressured to squeeze every last dollar, which pushes them past good UX and brand lines. Amol triages any controversial test into two buckets. Bucket one: so controversial you wouldn't ship it regardless (values, brand, or a hard red line like safety) — so don't even run it, the results don't matter. Bucket two: you don't love it but it's not a red line — run it, but demand a level of return proportional to how much it makes you cringe. Underpinning it all is a principle: be comfortable leaving money on the table.

Origin

Amol Avasare's decision framework as Anthropic's growth lead, rooted in Anthropic's safety-first mission and public-benefit-corporation structure; generalized from a broader life principle about not squeezing the last dollar.

Core principles

  • 01Every company's bucket-one red lines are different; name yours explicitly (for Anthropic, AI safety is bucket one)
  • 02If you'd never ship it, don't run the test — the results are irrelevant and running it wastes trust
  • 03For bucket-two tests, required return scales with the cringe factor: higher yikes demands higher payoff
  • 04Being comfortable leaving money on the table is a competitive advantage, not a weakness
  • 05The best products all forgo short-term metric impact to protect brand, quality, and user experience — which drives more growth long-term

How to run it

  1. 1

    Define your bucket-one red lines up front

    Decide which categories of test you will never ship regardless of results — the intersection of brand, values, and customer-friendliness. For Anthropic, anything compromising AI safety sits here.

    Pro tip Naming red lines in advance stops you from litigating them under the pressure of a promising hypothesis.

  2. 2

    Kill bucket-one tests before running them

    If a test is so controversial you wouldn't ship the winning variant anyway, don't run it at all. The results don't matter because the outcome is predetermined.

    Watch out Running a test you'd never ship burns engineering time and signals to the team that your red lines are negotiable.

  3. 3

    For bucket-two tests, gate on conviction and proportional return

    For tests you don't love but that aren't red lines, let someone with a strong, well-argued hypothesis run it — but require a higher expected return the more the test makes you cringe.

    Pro tip Calibrate: mild yikes needs modest upside; high cringe needs a high level of return to justify it.

  4. 4

    Stay comfortable leaving money on the table

    Accept forgoing metric impact to protect safety, brand, quality, and user experience. Zoom out past the quarter's numbers — the very best products all operate this way, and it compounds into more growth long-term.

    Pro tip Just as a founder shouldn't squeeze investors for the last dollar so they come back next round, don't squeeze users for the last metric.

    Watch out Squeezing every last dollar is one of the biggest mistakes hardcore growth practitioners make — it costs the repeat relationship.

In the wild

Anthropic withholding Claude for safety

Anthropic had a chatbot before ChatGPT launched but chose not to release it for safety reasons — the team didn't want to kick off an AI arms race. Time and again the company has accepted a significant commercial hit to prioritize safety.

The safety stance is treated as a bucket-one red line and, Amol argues, is becoming a significant long-term competitive advantage as stakes rise.

Common mistakes

Trying to squeeze every last dollar

Amol calls this one of the biggest mistakes hardcore growth practitioners make; extracting maximum short-term metric impact sacrifices brand, quality, and the repeat relationship that drives more growth long-term.

Running a test you'd never ship

If a variant sits in bucket one, its result can't change your decision, so running it wastes resource and erodes the credibility of your red lines — decide before, not after.

Is it for you?

Best for

Growth and product leaders who face pressure to optimize aggressively and need a repeatable way to weigh risky experiments against brand, safety, and values

Not ideal for

Pure performance-marketing contexts with no brand or values constraints, where the only variable is short-term ROI

From the transcript

One is when that test is so controversial that you you just should not run the thing because the results don't matter

1:14:30

if it's like a high level of like cringe or yikes then I want to see a high level of return

1:15:00

you usually need to be okay leaving money on the table

1:16:00

From the episode

Head of Growth (Anthropic): “Claude is growing itself at this point”

Amol Avasare