LLenny's Podcast
← All frameworks
Strategy

The Eval ROI Decision

Build evals where failure is catastrophic or you must win; vibe-check the rest.

Difficulty
Moderate
Time to result
~weeks to results
Steps
3
Confidence
90%

A pragmatic decision rule for when to invest in building evaluations for an AI feature versus shipping on vibes. Instead of treating evals as mandatory everywhere, you weigh the engineering cost against expected gain and the consequences of failure. Use it to decide where eval effort earns its keep.

Origin

Huyen answers Lenny's question about the online debate over whether AI products need evals. She lays out the executive's calculus: two engineers on evals might lift a metric from 80% to 82-85%, or those same engineers could launch a new feature worth far more.

Core principles

  • 01Evals are an investment that competes with feature work for the same engineers.
  • 02Consequence of failure sets the required rigor, not a blanket rule.
  • 03Good-enough-and-consistent can win; you don't have to be perfect at everything.

How to run it

  1. 1

    Estimate eval cost

    Ask how much effort building the eval takes (e.g., two engineers for some period).

  2. 2

    Estimate expected gain

    Ask how much the eval would realistically improve the metric. If it's an 80%-to-82% incremental bump, weigh that against the alternative use of those engineers.

    Pro tip Compare against the counterfactual: what could those engineers ship instead?

  3. 3

    Weight by failure severity

    If you operate at scale where failures are catastrophic, or the feature is a core competitive advantage, be tyrannical about evals. If it's a low-stakes helper, a vibe-check may be enough.

    Pro tip Reserve maximum eval rigor for user-facing, high-consequence, or must-win surfaces.

    Watch out No clear metric means poor visibility; the app can do something very dumb or costly without you knowing.

In the wild

The good-enough feature

A use case is already shipping, traffic is climbing, customers seem happy, but there's no exact metric. An engineer wants evals. The executive weighs two engineers on evals for a marginal lift against launching a new feature that yields much more.

For low-stakes features, the team ships and monitors; eval effort is redirected to higher-return work.

Common mistakes

Eval-everything zealotry

Insisting on rigorous evals for every minor feature drains engineers from higher-return work.

Vibes at catastrophic scale

Skipping evals where failures have severe consequences leaves you blind to expensive or dangerous failure modes.

Is it for you?

Best for

AI product leaders allocating scarce engineering time across features and quality infrastructure.

Not ideal for

Safety-critical or regulated systems where evals are non-negotiable regardless of ROI math.

From the transcript

it's all about like the question of like return investment

27:00

If you have if you operate at scale and where like failures can have like catastrophic consequences then you do need to be very tyrannical…

26:00

From the episode

Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)