Ambitious Retry (Pass@N) Tool Use
Get more from AI tools by asking for the ambitious change and fully restarting on failure instead of hammering the same attempt
- Difficulty
- Easy
- Time to result
- ~days to results
- Steps
- 4
- Confidence
- 87%
The gap between people who get huge value from AI tools and those who don't is not skill with the old workflow, it's willingness to ask for the big, ambitious change and to fully retry when it fails. Because models are stochastic, a completely fresh attempt succeeds far more often than repeatedly patching the same failed output. Treat the tool as pass@N, not pass@1.
Origin
Ben Mann's advice for future-proofing your career against AI, grounded in how model cards report pass@1 versus pass@N metrics, applied to everyday use of tools like Claude Code.
Core principles
- 01People who use new tools as if they were old tools tend to not succeed
- 02Ask for the ambitious outcome, not the incremental one the old workflow trained you to expect
- 03Models are stochastic: the same prompt can succeed or fail on identical inputs
- 04A clean restart beats banging on a broken attempt
How to run it
- 1
Ask for the ambitious change
Request the full, end-to-end outcome you actually want, not the small autocomplete-sized step you'd ask an old tool for.
Pro tip If you catch yourself scoping down out of habit, scope back up; the ceiling is higher than it feels.
Watch out Using the tool timidly, as if it were autocomplete, is the single biggest reason people underperform with it.
- 2
If it fails, start over completely
When the first attempt doesn't work, discard it and re-run from scratch rather than iterating on the broken output.
Pro tip Success rate on a full restart is 'much much higher' than continuing to patch a failed attempt.
- 3
Retry several times
Re-run the same request a few times; stochasticity means one of the fresh attempts often lands where earlier ones didn't.
Pro tip The 'dumbest' version is literally re-asking the identical prompt; it works surprisingly often.
- 4
Optionally, tell it what already failed
For a smarter retry, tell the model what you already tried and that it didn't work, so it explores a different path.
Pro tip Naming the dead ends steers it away from repeating them.
In the wild
Mann reports Anthropic's own legal and finance teams get a ton of value from Claude Code, using it to redline documents and run BigQuery analyses of customers and revenue metrics, despite it not being a traditional coding audience.
→ Non-engineering teams extracting engineering-grade leverage by being willing to jump into an unfamiliar, ambitious tool.
Common mistakes
Using new tools as if they were old tools
Treating an agent like autocomplete or basic chat caps you at the old workflow's ceiling and is why people fail to succeed with it.
Banging on the same failed attempt
Repeatedly patching a broken output has a much lower success rate than a clean restart, because the model is stochastic and a fresh run may simply get it right.
Is it for you?
Best for
Knowledge workers and non-engineers adopting agentic AI tools who want to future-proof their output against replacement
Not ideal for
Deterministic, safety-critical tasks where re-rolling a stochastic output until it looks right is unacceptable without verification
From the transcript
“People who use the new tools as if they were old tools tend to not succeed”
“are they asking for the ambitious change? And if it doesn't work the first time, asking three more times because our success rate when you…”
“These things are stochcastic and sometimes they'll figure it out and sometimes they won't”
From the episode
Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night
Ben Mann