The Determinism Test for AI Verticals
Favor AI domains where outputs can be tested quickly and objectively.
- Difficulty
- Advanced
- Time to result
- ~weeks to results
- Steps
- 6
- Confidence
- 91%
Simons evaluates AI verticals by asking how deterministic their feedback loops are. In software, generated code can be executed: it runs or it does not, and visual or functional tests can add further objective signals. Large numbers of application variants can be generated, tested, and used for reinforcement learning. That structure helps coding models improve rapidly and makes the resulting capability commercially legible. Other domains can be less deterministic. A legal outcome, for example, may depend on a judge, jury, politics, and the social standards of a particular time, making synthetic labels less reliable. The framework does not claim that only deterministic domains matter. It is a decision test: all else equal, favor opportunities where outputs can be evaluated quickly, cheaply, objectively, and across many permutations, especially when better performance unlocks a large existing market.
Origin
After Sonnet made Bolt viable, Simons asked why coding had advanced so quickly. He concluded that software gives model developers unusually objective, scalable feedback and used that mechanism—not a generic AGI forecast—to justify investing heavily in AI coding.
Core principles
- 01Fast objective feedback accelerates model improvement.
- 02Deterministic outputs create cleaner training and evaluation data.
- 03Tasks that can be permuted at scale offer more learning opportunities.
- 04Socially contingent outcomes are harder to label reliably.
- 05Technical mechanisms are stronger market evidence than generalized AI hype.
How to run it
- 1
Specify the output
Describe exactly what the model must produce for the vertical. Avoid broad labels such as 'legal AI' or 'health AI' that hide multiple tasks with different feedback properties.
Pro tip Score individual workflows, not entire industries.
- 2
Find the truth signal
Identify how the output can be judged correct or useful. Prefer signals generated by the task itself, such as executable tests, over opinions assembled after the fact.
Pro tip List both binary checks and graded checks; code can compile yet still produce a poor experience.
Watch out A proxy metric is not objective merely because it is numeric.
- 3
Price the feedback loop
Estimate the time, cost, and human labor required to generate one trustworthy evaluation. Faster and cheaper cycles support more learning.
- 4
Test permutation scale
Ask whether the task can be varied thousands of times without losing label quality. A broad space of automatically testable cases creates more reinforcement opportunities.
Pro tip Include edge cases and adversarial variants, not just common happy paths.
Watch out Synthetic variety is useless if every test repeats the same hidden bias.
- 5
Discount social contingency
Identify where correctness changes with a judge, audience, institution, politics, or time. Apply a higher uncertainty discount when those factors dominate the outcome.
Watch out Do not force inherently judgment-based work into a false binary.
- 6
Connect capability to value
Finally, verify that better model performance removes meaningful cost, delay, or access barriers in a large enough market. Determinism accelerates capability but does not guarantee a business.
Pro tip Look for users already paying humans or software to complete the task.
In the wild
Generated software can be compiled, run, and tested repeatedly, producing immediate evidence about whether it works. A predicted legal ruling can change with the judge, jury, politics, and prevailing social standards, so fabricated cases cannot provide equally reliable outcome labels. Simons uses this contrast to explain why coding models can improve through large-scale reinforcement loops.
→ Coding receives a higher determinism score and supports a stronger technical thesis for rapid capability gains.
A founder compares two ideas: generating database migrations that can be run against automated tests, and predicting whether executives will like a strategy memo. The first produces quick, repeatable correctness signals; the second depends heavily on individual taste and context.
→ The founder prioritizes the migration workflow for an initial product while treating memo evaluation as a human-in-the-loop feature.
Common mistakes
Scoring an industry instead of a task
A domain contains workflows with different feedback loops. Broad labels conceal whether the actual output can be tested.
Confusing executable with valuable
A task can be perfectly testable and commercially irrelevant. The final step must connect capability improvement to real economic value.
Ignoring graded quality
Code that runs can still be ugly, insecure, or unusable. Deterministic binary tests should be supplemented with functional and quality evaluations.
Is it for you?
Best for
Founders, investors, and product leaders evaluating where model capability is likely to improve rapidly.
Not ideal for
Domains where human judgment is itself the product and objective correctness is neither possible nor desirable.
From the episode
Inside Bolt: From near-death to ~$40m ARR in 5 months—one of the fastest-growing products in history
Eric Simons (founder and CEO of StackBlitz)