Value-Chain Eval
Before applying AI to your business, build a systematic test that measures how well it automates your core value chain
- Difficulty
- Moderate
- Time to result
- ~weeks to results
- Steps
- 3
- Confidence
- 90%
Foody's prescription for enterprises entering the 'era of evals': the prerequisite to deploying AI across a business is not the deployment itself but a repeatable way to measure success on the company's core value chain. If the model is the product, the eval is the product requirement document — and later the sales collateral proving the capability works. Each company has its own value chain (or a handful for multi-product firms) and must define the measurement for each.
Origin
Brendan Foody, CEO of Mercor, generalizing what he observed across working with 'six out of the magnificent seven' and 'all of the top five AI labs.' He credits Sarah Guo for the framing 'evals equal your new marketing' and notes Greg Brockman's 'evals are all you need.'
Core principles
- 01If the model is the product, the eval is the PRD
- 02Reinforcement learning is now effective enough that once a lab has an eval, it can 'hill climb' it — so the measure is the bottleneck, not the modeling
- 03The barrier to automating any workflow is measuring success, not the AI capability itself
- 04Evals double as sales collateral — the way you demonstrate a capability's efficacy
- 05Each company has its own value chain; a multi-product company has several
How to run it
- 1
Isolate your core value chain
Name the concrete deliverable your company produces for its end customer — e.g. an architecture firm produces architecture diagrams for its client. For multi-product companies, enumerate the handful of distinct value chains.
Pro tip Pick the output a customer actually pays for, not an internal process step.
- 2
Build a systematic measure of AI performance on that chain
Construct a test or rubric that scores how well AI reproduces the deliverable — the same way a professor writes a rubric for a graded deliverable, awarding points for each key element the output must contain.
Pro tip Have a domain expert author the rubric so it captures the mistakes and missed points a non-expert would never notice.
Watch out Vague success criteria produce un-hill-climbable evals; the rubric must be specific enough to score plus/minus points per criterion.
- 3
Use the eval to both measure progress and demonstrate capability
Run the eval as the benchmark you track improvement against, and reuse it externally as proof — the demonstration that your model or product actually automates the workflow customers care about.
Pro tip Move away from generic academic evals (GPQA, Olympiad math) toward the practical capabilities buyers care about.
Watch out Some enterprises avoid evaluating their business because the eval would expose that their value chain is being automated — that fear leaves the capability un-measured and un-improved.
In the wild
Foody's illustration: an architecture firm produces architecture diagrams for its end customer. The eval question is 'how can they effectively measure that' — building a systematic score for how well AI reproduces the diagram deliverable.
→ A measurable target that makes AI deployment across the business tractable.
A model writing a contract redline like a lawyer misses key points. A lawyer authors a rubric — 'plus however much of it identifies this or XYZ key point' — that scores the model's output the way a professor grades a deliverable.
→ A foundation for measuring model progress and generating reward signal to improve it.
Common mistakes
Treating evals as academic benchmarks
Pointing at GPQA or Olympiad math instead of the economically valuable capability your customers actually pay for. The useful eval measures your specific value chain.
Refusing to eval out of fear
Fortune 500 firms that won't evaluate their business because it would prove their value chain is automatable end up unable to improve or apply AI at all.
Is it for you?
Best for
Enterprise product leaders and operators deciding how to deploy AI across a company's core workflows
Not ideal for
Pure research settings chasing general intelligence rather than a specific business deliverable
From the transcript
“If the model is the product, then the eval is the product requirement document.”
“the core way to think about it is how can they build a test or systematic way to measure how well AI automates their core…”
“have a lawyer create a rubric similar to how a professor might create a rubric to create a deliverable for what are the things we…”
“and each company has its own value chain uh or maybe a handful of them if it's a multi-product company”
From the episode
Why experts writing AI evals is creating the fastest-growing companies in history
Brendan Foody (CEO of Mercor)