LLenny's Podcast
← All frameworks
InnovationAmol Avasare

The Automatable Growth Loop (CACHE)

Break growth experimentation into four evaluable stages an AI can hill-climb, keeping humans on alignment

Difficulty
Advanced
Time to result
~months to results
Steps
5
Confidence
85%

Anthropic's growth-platform team runs 'CACHE' (Claude Accelerates Sustainable Hypergrowth) to automate growth experimentation. The insight is to decompose the shipping life cycle into four discrete stages — identify opportunities, build, test against quality/brand bars, analyze learnings — each of which can be given an eval and scored so the model hill-climbs on it independently. A fifth stage, cross-functional stakeholder management, deliberately stays human. Progress is measured week-over-week: are results improving and human time decreasing in each stage?

Origin

Driven at Anthropic by Alexey Komissarouk (who teaches growth engineering at Reforge); became viable only with Opus 4.5/4.6 — Amol says it 'wasn't really possible' before.

Core principles

  • 01Decompose a workflow into discrete stages so each can be evaluated and improved on its own
  • 02Each stage needs an eval and a score, so you can hill-climb rather than judge by vibes
  • 03Encode judgment (brand guidelines, do's and don'ts) into a skill so the model can self-check, shrinking human review over time
  • 04Keep a human in the loop where the bottleneck is cross-functional alignment, not analysis
  • 05Measure the initiative's health by week-over-week improvement in results and reduction in human time per stage
  • 06Start on small, low-risk changes (copy, minor UI) where the model already performs, then expand scope as capability rises

How to run it

  1. 1

    Identify opportunities

    Have the model surface experiment ideas based on current and historical trends it has seen. Eval how good it is at spotting worthwhile opportunities versus noise.

  2. 2

    Build the feature ship-ready

    Have the model build the actual change and get it ready to ship. Score how reliably it produces something deployable.

  3. 3

    Test against quality and brand bars

    Check the change against your quality bar and brand bar before it goes live. Encode brand do's and don'ts into a skill the model consults so this gate can increasingly run without human review.

    Pro tip A skill loaded with brand guidelines, vision, mission, and goals lets the model reject off-brand ideas itself — and you can always unship a bad call.

    Watch out Brand is exactly where teams assume humans are needed forever; test whether a well-fed skill can do it instead.

  4. 4

    Analyze the data and gather learnings

    Once shipped, have the model analyze results and extract learnings that feed the next loop. Score the quality of the insights it returns.

  5. 5

    Keep humans on cross-functional alignment

    Do not automate stakeholder management. The one stage that stays human — especially for larger projects — is getting people aligned. For small changes you can skip it; for big ones it's the enduring human bottleneck.

    Pro tip Track the whole system by asking weekly: are results better, and is human time lower, in each stage? If yes, it's scaling.

    Watch out 'You will have AGI and it will still be impossible to get six people in a room to align' — don't expect automation to solve the alignment stage soon.

In the wild

Copy and minor UI experiments on autopilot

Anthropic kicked off CACHE only a couple of months before the interview, running it at small scale on copy changes and minor UI tweaks with a human still approving. It became viable only around Opus 4.5-4.6.

It 'prints money' at a win rate Amol compares to a junior PM two-to-three years in — not yet senior-PM level, but improving rapidly on the exponential.

Common mistakes

Trying to automate the whole workflow as one blob

Without decomposing into evaluable stages, you can't score or hill-climb — you're left judging a single opaque output by vibes instead of measuring which stage is weak.

Automating cross-functional alignment prematurely

Stakeholder management is the stage that still needs human brains; skipping it on large projects causes teams to spin their wheels and do overlapping work, negating the automation's gains.

Is it for you?

Best for

Growth and platform teams with high volumes of small, data-driven experiments who can define evals and want to scale experimentation with AI in the loop

Not ideal for

Large, controversial, or high-stakes bets dominated by cross-functional coordination, or teams without the eval discipline to score each stage

From the transcript

How can we use Claude to automate growth experimentation?

00:30

there's there's sort of four parts to it. One is there's um, identifying opportunities.

35:30

It ultimately prints money where I'd say that the win rate is like I would expect to see a junior PM to do better.

36:30

From the episode

Head of Growth (Anthropic): “Claude is growing itself at this point”

Amol Avasare