"Evals" Now Means Five Different Things (Semantic Diffusion)
The guest argues the word "evals" has been butchered into meaning error analysis, expert notes, LLM judges, feedback loops, and even public benchmarks. Invoking Martin Fowler's term "semantic diffusion," they warn that checking LM Arena is not doing evals. The real, agreed-upon goal is building an actionable feedback loop.
- Data labelers, PMs, and researchers all mean different things by "evals"
- Martin Fowler's "semantic diffusion": a term gets diluted until its meaning is lost
- Reading LM Arena or benchmarks is not the same as building your own eval dataset
- Whether you need an LLM judge depends entirely on context; don't over-prescribe
“someone comes up with a term everybody starts butchering it with their own definitions and then you kind of lose the actual definition of it.”
“You're not doing eval. That's not eval. Those are model.”