LLenny's Podcast
← All frameworks
StrategyEdwin Chen (Surge AI)

Deep Human Evaluation Over Leaderboards

Measure real progress with experts who work through the answer, not crowds who vibe for two seconds.

Difficulty
Advanced
Time to result
~weeks to results
Steps
4
Confidence
87%

A method for measuring the true quality of a system when the popular metrics are gameable. Instead of trusting benchmarks or crowd-voted leaderboards, you put domain experts into realistic role-based scenarios and have them work deeply through every output — checking the code, re-deriving the equations — across many dimensions. Use it whenever you suspect your headline metric is being optimized instead of the thing it's supposed to represent.

Origin

Chen distrusts AI benchmarks and leaderboards like LM Arena because voters skim for two seconds and pick whatever looks flashiest, letting a fully-hallucinated answer win on emojis and length. Surge measures model progress instead through deep human evaluations by expert annotators.

Core principles

  • 01Popular benchmarks are often literally wrong and are easy to hill-climb via gaming.
  • 02Crowd voting rewards superficial polish (emojis, length, markdown) over accuracy.
  • 03Real evaluation requires experts who actually do the task, not skimmers who vibe.
  • 04Evaluate across many dimensions — accuracy, instruction-following, correctness — not one gut impression.

How to run it

  1. 1

    Recruit genuine domain experts

    Use evaluators who are at the top of their field for the domain being tested — a Nobel-level physicist, a working coder, a teacher.

    Watch out Casual users skim; they won't catch the errors that matter.

  2. 2

    Put them in realistic role scenarios

    Have each expert converse with the system as they actually would in their work — pushing their own research frontier, building lesson plans, solving daily engineering problems.

    Pro tip Span many topics and roles so you see performance across the real distribution of use.

  3. 3

    Make them work through, not skim

    Require evaluators to actually run the code, double-check the equations, and verify claims — evaluating deeply rather than judging on appearance.

    Pro tip Score explicit dimensions like accuracy and instruction-following that casual users ignore.

    Watch out A flashy, emoji-laden, long answer can be entirely hallucinated — appearance is the trap.

  4. 4

    Distrust the gameable metric

    Treat benchmark and leaderboard scores as marketing signals that get gamed via prompt tweaks, run counts, and leakage — not as ground truth.

    Watch out Optimizing for the benchmark instead of the real world climbs the benchmark anyway — that's just another form of gaming.

In the wild

LM Arena vs deep annotation

On LM Arena, random voters skim two responses for two seconds and pick the flashiest — so the fastest way to climb is more emojis, bolding, and tripled length, even if the model hallucinates and gets the answer wrong. Surge's expert annotators instead read closely and fact-check every dimension.

Chen argues the models with the best leaderboard scores are often the worst on real accuracy and instruction-following.

IMO gold vs parsing a PDF

Models win IMO gold medals yet still struggle to parse PDFs, because contest math has clean objective answers that are easy to hill-climb, while real-world messiness does not.

Illustrates why benchmark performance is weakly correlated with real capability.

Common mistakes

Trusting the crowd AB test

When a pop-up asks you to pick the better response, you vibe and pick what looks slickest — that's not evaluation, and optimizing for it degrades the product.

Assuming benchmarks are correct

Even researchers don't realize many benchmarks contain wrong answers and messiness, so a high score can be measuring the wrong thing.

Is it for you?

Best for

Teams building or buying AI/ML systems (or any product) where the public metric can be gamed and real quality is subtle.

Not ideal for

Low-stakes tasks where a cheap approximate metric is good enough and expert time isn't worth it.

From the transcript

they are not just skimming the responses, they're actually working through the responses deeply themselves.

21:00

They're just vibing and picking whatever response looks slashiest.

21:30

The easiest way to climb Alam Marina, it's adding crazy voting. It's doubling the number of emojis. It's tripling the length of your model responses.

24:00

From the episode

The 100-person AI lab that became Anthropic and Google's secret weapon

Edwin Chen (Surge AI)