LLenny's Podcast
← All frameworks
StrategyJason Droege

The AI-Fit Accuracy Test

Deploy AI where the human baseline is low-accuracy — not where it's already near-perfect and you need the last 2%

Difficulty
Moderate
Time to result
~months to results
Steps
3
Confidence
85%

Droege's heuristic for where AI actually delivers value: measure the current human process's accuracy. If humans are only 10-20% accurate or satisfied, AI jumping to 50-80% is pure gain; if the process is already 98% accurate, expecting AI to close the final 2% is 'not totally there yet.' The remaining gap always needs a human-in-the-loop escalation path.

Origin

Jason Droege's framing from Scale AI's enterprise deployments (healthcare document triage, insurance claims management), distinguishing where 'good' beats 'correct' in probabilistic systems.

Core principles

  • 01AI wins biggest where the human baseline is low-accuracy or low-satisfaction
  • 02Closing the last 1-2% of an already-excellent process is where AI still falls short — like adding 'nines' of uptime, each one costs an order of magnitude
  • 03These are probabilistic systems, so the target is 'what does good look like,' not a single correct answer
  • 04Any deployment must route the low-confidence remainder to a human for guidance
  • 05Robust automation of an important process realistically takes six to twelve months, not minutes

How to run it

  1. 1

    Measure the human baseline

    Quantify how accurate or satisfying the current human process is before deciding whether AI fits. This number determines the size of the available win.

    Watch out Don't be seduced by demo videos that imply minutes-to-value; real robustness takes six to twelve months.

  2. 2

    Match AI to low-baseline processes

    Target processes sitting at ~10-20% accuracy where AI lifting them to 50-80% creates obvious net value and everyone is 'in the green.' Avoid staking success on AI closing the final 2% of a 98%-accurate process.

    Pro tip Frame the goal as 'what does good look like,' since these are probabilistic systems delivering best-recommendation judgments, not guaranteed-correct answers.

    Watch out Every additional 'nine' of reliability costs roughly an order of magnitude of investment — pricing the last 2% like the first 60% will burn you.

  3. 3

    Engineer the human-in-the-loop remainder

    Build the system so that when confidence is low, it escalates to a human for feedback and guidance rather than acting. The humans stay net-positive contributors on the residual decisions.

    Pro tip Establish evals — a comprehensive benchmark of 'what good looks like' — so the system knows when it's below bar and must escalate.

    Watch out Mission-critical agentic workflows demanding very high accuracy are exactly where the residual-handling matters most.

In the wild

Healthcare rare-case document triage

Specialist doctors faced 200-300 pages of mixed-format documentation per rare case and could only skim it. Scale built a tool to read the document and surface the top 5-10 things to consider — the human baseline (a rushed skim) was low, so AI added clear value, and it once flagged a non-obvious allergy that conflicted with a planned medication.

Improved productivity against a huge backlog and caught a correlation 'that would have even been hard for a human being to do' — a high-value win precisely because the human baseline was weak.

Insurance claims management plateau

In claims workflows the systems reach ~60-70% of the way and the human mind assumes the rest is 'no big deal' — but like adding nines of data-center uptime, each increment of reliability is an order-of-magnitude harder, so full automation is far costlier than it looks.

Reframes expectations: robust automation of an important process takes six to twelve months, not the minutes implied by hype.

Common mistakes

Expecting AI to close the last 2%

Pointing AI at a 98%-accurate human process and expecting it to nail the remaining 2% overreaches current capability — 'not totally there yet' — and each additional nine of reliability costs an order of magnitude.

Believing the minutes-to-value demos

Pilots fail partly because it's trivially easy to spin one up, and demo videos imply instant results; truly robust automation takes six to twelve months of legal, policy, regulatory, and change-management work.

Is it for you?

Best for

Enterprise leaders and product teams deciding which workflows to automate with AI first

Not ideal for

Already near-perfect processes where the only remaining gain is the final fraction of a percent

From the transcript

if you have a human process that is like 10 or 20% AC like accurate or 10 or 20% liked AI is awesome because it…

32:30

if you have uh a human process, a workflow that is 98% accurate and you expect an AI system to get you the remainder of…

33:00

these things take six to 12 months to get them truly, you know, robust enough where like an important process can be automated

39:30

From the episode

First interview with Scale AI’s CEO: $14B Meta deal, what’s working in enterprise AI, and what frontier labs are building next

Jason Droege