LLenny's Podcast
← All frameworks
ProductivityCat Wu (Head of Product, Claude Code)

Three-Legged Model Evaluation

Judge a model-harness combo with heavy usage, a trusted five-person taste panel, and ~10 sharp evals

Difficulty
Moderate
Time to result
~weeks to results
Steps
3
Confidence
90%

To build the rare skill of judging what a model can and can't do, Cat leans on three complementary sources of signal: deep personal usage of the model, a small handpicked group of people whose feedback is genuinely qualified, and a modest set of well-crafted evals. The key insight is that not everyone's feedback is equally valuable, and you don't need hundreds of evals — ten great ones quantify the goal and progress toward it.

Origin

Cat Wu's method for evaluating Claude and Claude Code models and harnesses.

Core principles

  • 01Spend a large share of your time actually talking to and using the model
  • 02Feedback quality varies wildly — identify the few people who articulate model quality accurately
  • 03Evals don't need to be numerous; ten great ones beat hundreds of mediocre ones
  • 04Evals quantify the goal, current progress, and what's missing — turning vibes into measurement

How to run it

  1. 1

    Log heavy hands-on time with the model

    Spend a substantial fraction of your time pushing the model's boundaries so you develop a strong sense of what it's not good at and why it makes the mistakes it does.

    Pro tip Cat spends roughly 30% of her time probing what the tools can and can't do.

  2. 2

    Assemble a trusted five-person taste panel

    Find the handful of people who are much better than others at articulating what makes a specific model or model-harness combination good, and rely on them for fast, high-signal feedback.

    Pro tip A quick round-the-table 'what's your vibe on the model?' at a team lunch surfaces hypotheses fast.

    Watch out Lots of people will give feedback, but not everyone's is qualified — weight the trusted few over volume.

  3. 3

    Build ~10 great evals, not hundreds

    Write a small set of high-quality evals that quantify what the goal is, how far along you are, and what you're missing. Use them to convert subjective impressions into measurable progress.

    Pro tip Jump into eval-writing specifically when a feature needs more product definition — the output is a handful of evals plus the prompt that raised the success rate.

In the wild

Team-lunch vibe checks

When a new model is being tested, the Claude team gathers feedback fast by going person-to-person at team lunches asking each one's 'vibe' — surfacing observations like the model being too abrupt, over-writing memories, or not testing itself enough. These become hypotheses the team then verifies against real data.

Rapid, qualified feedback that directs which data patterns to investigate and validate.

Common mistakes

Weighting all feedback equally

Treating every reviewer's opinion the same drowns out the few people who can actually articulate model quality, slowing and muddying decisions.

Believing you need hundreds of evals

The perceived burden of building large eval suites stops teams from building any. Ten great evals already quantify the goal and progress — the bar to start is low.

Is it for you?

Best for

AI PMs and engineers who need to evaluate model or harness quality and want signal beyond gut feel

Not ideal for

Features where success is trivially verifiable and formal evals add no product definition

From the transcript

finding a group of those like five people you trust is really important for getting very fast feedback.

54:30

You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify…

55:00

From the episode

How Anthropic’s product team moves faster than anyone else

Cat Wu (Head of Product, Claude Code)