Three-Legged Model Evaluation
Judge a model-harness combo with heavy usage, a trusted five-person taste panel, and ~10 sharp evals
- Difficulty
- Moderate
- Time to result
- ~weeks to results
- Steps
- 3
- Confidence
- 90%
To build the rare skill of judging what a model can and can't do, Cat leans on three complementary sources of signal: deep personal usage of the model, a small handpicked group of people whose feedback is genuinely qualified, and a modest set of well-crafted evals. The key insight is that not everyone's feedback is equally valuable, and you don't need hundreds of evals — ten great ones quantify the goal and progress toward it.
Origin
Cat Wu's method for evaluating Claude and Claude Code models and harnesses.
Core principles
- 01Spend a large share of your time actually talking to and using the model
- 02Feedback quality varies wildly — identify the few people who articulate model quality accurately
- 03Evals don't need to be numerous; ten great ones beat hundreds of mediocre ones
- 04Evals quantify the goal, current progress, and what's missing — turning vibes into measurement
How to run it
- 1
Log heavy hands-on time with the model
Spend a substantial fraction of your time pushing the model's boundaries so you develop a strong sense of what it's not good at and why it makes the mistakes it does.
Pro tip Cat spends roughly 30% of her time probing what the tools can and can't do.
- 2
Assemble a trusted five-person taste panel
Find the handful of people who are much better than others at articulating what makes a specific model or model-harness combination good, and rely on them for fast, high-signal feedback.
Pro tip A quick round-the-table 'what's your vibe on the model?' at a team lunch surfaces hypotheses fast.
Watch out Lots of people will give feedback, but not everyone's is qualified — weight the trusted few over volume.
- 3
Build ~10 great evals, not hundreds
Write a small set of high-quality evals that quantify what the goal is, how far along you are, and what you're missing. Use them to convert subjective impressions into measurable progress.
Pro tip Jump into eval-writing specifically when a feature needs more product definition — the output is a handful of evals plus the prompt that raised the success rate.
In the wild
When a new model is being tested, the Claude team gathers feedback fast by going person-to-person at team lunches asking each one's 'vibe' — surfacing observations like the model being too abrupt, over-writing memories, or not testing itself enough. These become hypotheses the team then verifies against real data.
→ Rapid, qualified feedback that directs which data patterns to investigate and validate.
Common mistakes
Weighting all feedback equally
Treating every reviewer's opinion the same drowns out the few people who can actually articulate model quality, slowing and muddying decisions.
Believing you need hundreds of evals
The perceived burden of building large eval suites stops teams from building any. Ten great evals already quantify the goal and progress — the bar to start is low.
Is it for you?
Best for
AI PMs and engineers who need to evaluate model or harness quality and want signal beyond gut feel
Not ideal for
Features where success is trivially verifiable and formal evals add no product definition
From the transcript
“finding a group of those like five people you trust is really important for getting very fast feedback.”
“You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify…”
From the episode
How Anthropic’s product team moves faster than anyone else
Cat Wu (Head of Product, Claude Code)