LLenny's Podcast
← All frameworks
CommunicationNick Turley (Head of ChatGPT at OpenAI)

Evals as Articulating Success

An eval is just a clear spec of ideal behavior — the shared language of AI product work

Difficulty
Easy
Time to result
~weeks to results
Steps
4
Confidence
85%

Demystifies evals for product people: an eval is nothing more than clearly articulating what success/ideal behavior looks like for each use case, in a form useful for improving the model. It's the timeless product wisdom of 'define success before you build' expressed in a new mechanism — and it's the lingua franca between product people and AI researchers. You can write one in a spreadsheet.

Origin

Turley says he was 'writing evals before I knew what an eval was' — just outlining clearly specified ideal behavior for use cases — until someone told him that's what an eval is. He realized the research world's evaluation benchmarks were the same idea, and that evals are how you communicate to researchers what the product should do.

Core principles

  • 01An eval is just articulating success in a way useful for training/improving the model
  • 02It's the same eternal wisdom as 'specify success before you build,' in a new context
  • 03Evals are the shared language between product and AI research
  • 04No special technical magic required — a spreadsheet works
  • 05Real-world failure cases are the highest-value eval material as benchmarks saturate

How to run it

  1. 1

    Specify ideal behavior per use case

    For each use case, write down clearly and concretely what the model should do — the ideal behavior — before building.

    Pro tip You don't need to know the term 'eval' to start; just outline the specified ideal behavior.

  2. 2

    Capture it in any medium

    Put it in a spreadsheet or anywhere. Don't treat it as technical magic requiring special tooling.

    Watch out Treating evals as an opaque research-only artifact keeps product people out of the loop that most needs them.

  3. 3

    Use it as the product/research handshake

    Hand the eval to research/ML teams as the precise articulation of what the product should be doing — the lingua franca that communicates intent.

  4. 4

    Feed real failure cases in

    As benchmarks saturate, harvest real-world scenarios where the product fails and turn them into evals so ML teams know exactly what to climb.

    Pro tip 'People are trying to do X and the model's failing in ways Y — now let's make those things really good.'

In the wild

Discovering evals by accident

Turley outlined ideal behavior for various use cases as a product exercise, then learned there was an entire research world of evaluation benchmarks describing the same thing — and that this was how to communicate product intent to AI researchers.

Evals became the shared vocabulary bridging product and research at OpenAI.

Common mistakes

Treating evals as technical magic

Turley actively wants to demystify the term — it's just articulating success, doable in a spreadsheet, not something only ML researchers can touch.

Relying on saturated benchmarks

As standard benchmarks saturate, they stop discriminating. Real-world failure cases are where the signal is, and only shipping produces them.

Is it for you?

Best for

Product managers and builders working with AI research/ML teams who need a shared language for desired model behavior

Not ideal for

Non-AI product work where success is already measured by conventional analytics

From the transcript

I started writing evals before I knew what an eval was because like I was just outlining sort of very clearly specified ideal behavior for…

1:14:30

it's really just about articulating success in a way that is maximally useful for for training bots

the best way to then go articulate to your team, especially your MEL teams, what to climb on

1:14:00

From the episode

Inside ChatGPT: The fastest-growing product in history

Nick Turley (Head of ChatGPT at OpenAI)