Evals as Articulating Success
An eval is just a clear spec of ideal behavior — the shared language of AI product work
- Difficulty
- Easy
- Time to result
- ~weeks to results
- Steps
- 4
- Confidence
- 85%
Demystifies evals for product people: an eval is nothing more than clearly articulating what success/ideal behavior looks like for each use case, in a form useful for improving the model. It's the timeless product wisdom of 'define success before you build' expressed in a new mechanism — and it's the lingua franca between product people and AI researchers. You can write one in a spreadsheet.
Origin
Turley says he was 'writing evals before I knew what an eval was' — just outlining clearly specified ideal behavior for use cases — until someone told him that's what an eval is. He realized the research world's evaluation benchmarks were the same idea, and that evals are how you communicate to researchers what the product should do.
Core principles
- 01An eval is just articulating success in a way useful for training/improving the model
- 02It's the same eternal wisdom as 'specify success before you build,' in a new context
- 03Evals are the shared language between product and AI research
- 04No special technical magic required — a spreadsheet works
- 05Real-world failure cases are the highest-value eval material as benchmarks saturate
How to run it
- 1
Specify ideal behavior per use case
For each use case, write down clearly and concretely what the model should do — the ideal behavior — before building.
Pro tip You don't need to know the term 'eval' to start; just outline the specified ideal behavior.
- 2
Capture it in any medium
Put it in a spreadsheet or anywhere. Don't treat it as technical magic requiring special tooling.
Watch out Treating evals as an opaque research-only artifact keeps product people out of the loop that most needs them.
- 3
Use it as the product/research handshake
Hand the eval to research/ML teams as the precise articulation of what the product should be doing — the lingua franca that communicates intent.
- 4
Feed real failure cases in
As benchmarks saturate, harvest real-world scenarios where the product fails and turn them into evals so ML teams know exactly what to climb.
Pro tip 'People are trying to do X and the model's failing in ways Y — now let's make those things really good.'
In the wild
Turley outlined ideal behavior for various use cases as a product exercise, then learned there was an entire research world of evaluation benchmarks describing the same thing — and that this was how to communicate product intent to AI researchers.
→ Evals became the shared vocabulary bridging product and research at OpenAI.
Common mistakes
Treating evals as technical magic
Turley actively wants to demystify the term — it's just articulating success, doable in a spreadsheet, not something only ML researchers can touch.
Relying on saturated benchmarks
As standard benchmarks saturate, they stop discriminating. Real-world failure cases are where the signal is, and only shipping produces them.
Is it for you?
Best for
Product managers and builders working with AI research/ML teams who need a shared language for desired model behavior
Not ideal for
Non-AI product work where success is already measured by conventional analytics
From the transcript
“I started writing evals before I knew what an eval was because like I was just outlining sort of very clearly specified ideal behavior for…”
“it's really just about articulating success in a way that is maximally useful for for training bots”
“the best way to then go articulate to your team, especially your MEL teams, what to climb on”
From the episode
Inside ChatGPT: The fastest-growing product in history
Nick Turley (Head of ChatGPT at OpenAI)