LLenny's Podcast
← All frameworks
ProductivityBen Mann

Ambitious Retry (Pass@N) Tool Use

Get more from AI tools by asking for the ambitious change and fully restarting on failure instead of hammering the same attempt

Difficulty
Easy
Time to result
~days to results
Steps
4
Confidence
87%

The gap between people who get huge value from AI tools and those who don't is not skill with the old workflow, it's willingness to ask for the big, ambitious change and to fully retry when it fails. Because models are stochastic, a completely fresh attempt succeeds far more often than repeatedly patching the same failed output. Treat the tool as pass@N, not pass@1.

Origin

Ben Mann's advice for future-proofing your career against AI, grounded in how model cards report pass@1 versus pass@N metrics, applied to everyday use of tools like Claude Code.

Core principles

  • 01People who use new tools as if they were old tools tend to not succeed
  • 02Ask for the ambitious outcome, not the incremental one the old workflow trained you to expect
  • 03Models are stochastic: the same prompt can succeed or fail on identical inputs
  • 04A clean restart beats banging on a broken attempt

How to run it

  1. 1

    Ask for the ambitious change

    Request the full, end-to-end outcome you actually want, not the small autocomplete-sized step you'd ask an old tool for.

    Pro tip If you catch yourself scoping down out of habit, scope back up; the ceiling is higher than it feels.

    Watch out Using the tool timidly, as if it were autocomplete, is the single biggest reason people underperform with it.

  2. 2

    If it fails, start over completely

    When the first attempt doesn't work, discard it and re-run from scratch rather than iterating on the broken output.

    Pro tip Success rate on a full restart is 'much much higher' than continuing to patch a failed attempt.

  3. 3

    Retry several times

    Re-run the same request a few times; stochasticity means one of the fresh attempts often lands where earlier ones didn't.

    Pro tip The 'dumbest' version is literally re-asking the identical prompt; it works surprisingly often.

  4. 4

    Optionally, tell it what already failed

    For a smarter retry, tell the model what you already tried and that it didn't work, so it explores a different path.

    Pro tip Naming the dead ends steers it away from repeating them.

In the wild

Legal and finance teams using Claude Code

Mann reports Anthropic's own legal and finance teams get a ton of value from Claude Code, using it to redline documents and run BigQuery analyses of customers and revenue metrics, despite it not being a traditional coding audience.

Non-engineering teams extracting engineering-grade leverage by being willing to jump into an unfamiliar, ambitious tool.

Common mistakes

Using new tools as if they were old tools

Treating an agent like autocomplete or basic chat caps you at the old workflow's ceiling and is why people fail to succeed with it.

Banging on the same failed attempt

Repeatedly patching a broken output has a much lower success rate than a clean restart, because the model is stochastic and a fresh run may simply get it right.

Is it for you?

Best for

Knowledge workers and non-engineers adopting agentic AI tools who want to future-proof their output against replacement

Not ideal for

Deterministic, safety-critical tasks where re-rolling a stochastic output until it looks right is unacceptable without verification

From the transcript

People who use the new tools as if they were old tools tend to not succeed

18:30

are they asking for the ambitious change? And if it doesn't work the first time, asking three more times because our success rate when you…

19:00

These things are stochcastic and sometimes they'll figure it out and sometimes they won't

20:30

From the episode

Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night

Ben Mann