✶Explainer04:00
The numbers behind AI coding at OpenAI: 95% of engineers, 100% of PRs
Sherwin shares concrete internal metrics on how deeply Codex is embedded in OpenAI's engineering. Nearly all engineers use it daily, every PR is reviewed by it, and the code that hits production is close to 100% AI-generated first. He also notes a widening productivity gap: heavy Codex users open far more PRs than light users.
- 95% of engineers use Codex on a daily basis
- 100% of PRs are reviewed by Codex daily before merge
- Close to 100% of code is generated by AI first, then reviewed
- Engineers who use Codex more open 70% more PRs, and the gap keeps widening
“So 95% of engineers um use codeex. Um 100% of our PRs are reviewed by codeex daily as well.”
“So uh they're actually opening 70% more PRs uh and uh than than the engineers who aren't using codecs as much.”
#codex#openai#engineering productivity#ai coding
✶Explainer15:00
How Codex turned code review from the worst job into a 2-minute task
Sherwin explains that the tasks engineers hand to AI first are the ones they hate most, which is why work is more fun now. Code review is a prime example: Codex reviews 100% of PRs and collapses review time. For small PRs, teams increasingly trust Codex as the second pair of eyes rather than requiring a human reviewer.
- The most-hated, most-boring tasks get handed to AI first
- Codex reviews 100% of PRs; review drops from 10-15 minutes to 2-3 minutes
- For small PRs, teams often skip human review and trust Codex
- CI, lint fixes, and deployment steps are heavily automated via Codex
“Yeah, I mean one thing is Codex reviews 100% of all of our PRs at this point.”
“it makes you know code reviews go from a you know I don't know 10 15 minute task to sometimes even just like a two…”
#code review#codex#automation#developer workflow
✶Explainer19:30
Why AI lets managers run far larger teams than 6-8 reports
Sherwin argues the manager role has changed less than the IC role so far, but the trend is clear. AI tools that surface organizational context, do research, and write deep-research performance reviews will let people managers operate at much higher leverage, breaking past the old best practice of six to eight direct reports.
- The manager role has changed less than the IC role, but trends point one way
- ChatGPT hooked to GitHub, Notion, and Google Docs speeds up performance reviews
- Managers will manage far more than the current 6-8 direct-report norm
- The same leverage applies to non-engineering functions like support and operations
“I think managers will be able to manage much larger teams in this world kind of like how you know like software engineers are managing…”
#management#org design#ai tools#leverage
✶Explainer37:30
Why most enterprise AI deployments have negative ROI (and the fix)
Sherwin isn't surprised many AI deployments are net-negative: Silicon Valley forgets it lives in a bubble, and most workers are basic users, not power users. The failure pattern is top-down mandates divorced from real work. The winning setup pairs top-down buy-in with bottoms-up adoption, often via a dedicated internal tiger team that evangelizes and builds best practices.
- Many enterprise AI deployments are likely negative ROI
- Silicon Valley lives in a bubble; most employees are basic, not power, users
- Winning deployments combine top-down buy-in with bottoms-up adoption
- Staff a dedicated 'tiger team' of excited, technical-adjacent people to spread best practices
- These evangelists are often not software engineers, but Excel-wizard operations types
“I think we in Silicon Valley just forget that we live in a bubble.”
“find or maybe even staff a full-time team internally that is this kind of tiger team internally”
#ai adoption#enterprise#roi#change management
✶Explainer50:00
Multi-hour coherent tasks and audio are what's next
Asked where the models are heading in the next 6-18 months, Sherwin points to two things. First, task length: the METR benchmark shows models trending toward multi-hour coherent software tasks, which will reshape the products built around them. Second, audio, which he calls hugely underrated because most of the world's business runs on talking, not text.
- The METR benchmark tracks how long a task models can do 50%/80% of the time
- Products today optimize for ~10-minute interactive tasks; multi-hour is coming
- In 12-18 months, models may run 6-hour dispatched tasks coherently
- Native speech-to-speech audio models are a hugely underrated enterprise domain
“in the next 12 to 18 months we could see models that could do multi-hour long tasks very very coherently.”
“A lot of the world's business is done via audio.”
#model roadmap#metr benchmark#audio#long-running agents