The AI-Fit Accuracy Test
Deploy AI where the human baseline is low-accuracy — not where it's already near-perfect and you need the last 2%
- Difficulty
- Moderate
- Time to result
- ~months to results
- Steps
- 3
- Confidence
- 85%
Droege's heuristic for where AI actually delivers value: measure the current human process's accuracy. If humans are only 10-20% accurate or satisfied, AI jumping to 50-80% is pure gain; if the process is already 98% accurate, expecting AI to close the final 2% is 'not totally there yet.' The remaining gap always needs a human-in-the-loop escalation path.
Origin
Jason Droege's framing from Scale AI's enterprise deployments (healthcare document triage, insurance claims management), distinguishing where 'good' beats 'correct' in probabilistic systems.
Core principles
- 01AI wins biggest where the human baseline is low-accuracy or low-satisfaction
- 02Closing the last 1-2% of an already-excellent process is where AI still falls short — like adding 'nines' of uptime, each one costs an order of magnitude
- 03These are probabilistic systems, so the target is 'what does good look like,' not a single correct answer
- 04Any deployment must route the low-confidence remainder to a human for guidance
- 05Robust automation of an important process realistically takes six to twelve months, not minutes
How to run it
- 1
Measure the human baseline
Quantify how accurate or satisfying the current human process is before deciding whether AI fits. This number determines the size of the available win.
Watch out Don't be seduced by demo videos that imply minutes-to-value; real robustness takes six to twelve months.
- 2
Match AI to low-baseline processes
Target processes sitting at ~10-20% accuracy where AI lifting them to 50-80% creates obvious net value and everyone is 'in the green.' Avoid staking success on AI closing the final 2% of a 98%-accurate process.
Pro tip Frame the goal as 'what does good look like,' since these are probabilistic systems delivering best-recommendation judgments, not guaranteed-correct answers.
Watch out Every additional 'nine' of reliability costs roughly an order of magnitude of investment — pricing the last 2% like the first 60% will burn you.
- 3
Engineer the human-in-the-loop remainder
Build the system so that when confidence is low, it escalates to a human for feedback and guidance rather than acting. The humans stay net-positive contributors on the residual decisions.
Pro tip Establish evals — a comprehensive benchmark of 'what good looks like' — so the system knows when it's below bar and must escalate.
Watch out Mission-critical agentic workflows demanding very high accuracy are exactly where the residual-handling matters most.
In the wild
Specialist doctors faced 200-300 pages of mixed-format documentation per rare case and could only skim it. Scale built a tool to read the document and surface the top 5-10 things to consider — the human baseline (a rushed skim) was low, so AI added clear value, and it once flagged a non-obvious allergy that conflicted with a planned medication.
→ Improved productivity against a huge backlog and caught a correlation 'that would have even been hard for a human being to do' — a high-value win precisely because the human baseline was weak.
In claims workflows the systems reach ~60-70% of the way and the human mind assumes the rest is 'no big deal' — but like adding nines of data-center uptime, each increment of reliability is an order-of-magnitude harder, so full automation is far costlier than it looks.
→ Reframes expectations: robust automation of an important process takes six to twelve months, not the minutes implied by hype.
Common mistakes
Expecting AI to close the last 2%
Pointing AI at a 98%-accurate human process and expecting it to nail the remaining 2% overreaches current capability — 'not totally there yet' — and each additional nine of reliability costs an order of magnitude.
Believing the minutes-to-value demos
Pilots fail partly because it's trivially easy to spin one up, and demo videos imply instant results; truly robust automation takes six to twelve months of legal, policy, regulatory, and change-management work.
Is it for you?
Best for
Enterprise leaders and product teams deciding which workflows to automate with AI first
Not ideal for
Already near-perfect processes where the only remaining gain is the final fraction of a percent
From the transcript
“if you have a human process that is like 10 or 20% AC like accurate or 10 or 20% liked AI is awesome because it…”
“if you have uh a human process, a workflow that is 98% accurate and you expect an AI system to get you the remainder of…”
“these things take six to 12 months to get them truly, you know, robust enough where like an important process can be automated”
From the episode
First interview with Scale AI’s CEO: $14B Meta deal, what’s working in enterprise AI, and what frontier labs are building next
Jason Droege