Randomized Tiered Trial for AI Productivity
Measure whether AI tools help by running a randomized trial split across performance tiers.
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 3
- Confidence
- 82%
A measurement method for the notoriously hard question of whether AI coding tools actually improve productivity. Rather than relying on handwavy 'it feels better' or bad proxies like lines of code, a manager privately segments the team by performance tier and randomly gives half of each tier access to the tool, then observes the difference. Use it before spending big on team-wide AI subscriptions.
Origin
Huyen recounts a friend's company that ran exactly this design on a ~30-40 person engineering team, splitting engineers into best/average/lowest performing buckets (kept private) and randomly granting half of each bucket Cursor access.
Core principles
- 01Productivity is hard to measure; 'more PRs' or lines of code are poor proxies.
- 02A randomized control isolates the tool's effect from selection bias.
- 03Effects differ by skill tier, so measure per-tier, not just in aggregate.
How to run it
- 1
Segment by performance tier (privately)
Split the team into buckets such as best, average, and lowest performing, without telling them the labels.
Pro tip Keeping the buckets private avoids morale effects and behavior change from being labeled.
- 2
Randomize access within each tier
Give half of each bucket access to the AI tool (e.g., Cursor) and withhold it from the other half.
- 3
Observe the tier-level difference over time
Compare productivity between the tool and no-tool halves within each tier to see who actually benefits.
Watch out Results vary widely by company; one firm found senior engineers gained most, another found seniors most resistant. Don't assume your result generalizes.
In the wild
A ~30-40 person team split into three private performance buckets and randomly gave half of each Cursor access. Over time the highest-performing engineers got the biggest boost, then average performers, with lowest performers benefiting least.
→ A defensible, per-tier read on where the AI tool actually helped, instead of a vibe.
Common mistakes
Lines-of-code as productivity
More PRs or more code generated is not a valid productivity metric and can mislead investment decisions.
Assuming one universal result
Because different companies see opposite tier effects, copying another org's conclusion without your own trial is risky.
Is it for you?
Best for
VPs and managers deciding whether expensive team-wide AI coding subscriptions are worth it.
Not ideal for
Very small teams where a randomized split has too little signal, or orgs unwilling to run controlled experiments.
From the transcript
“The main thing is like it's really hard to measure productivity”
“there's a randomized trial so like they give like half of each of each group like access to like to like cursor”
From the episode
Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)