LLenny's Podcast
← All episodes
Alexander Embiricos (OpenAI Codex Product Lead)14 December 2025

Why humans are AI’s biggest bottleneck (and what’s coming in 2026)

6Frameworks
15Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Hot Take· 4

Hot Take14:00

AI Products Are Actually Hard to Use — And Proactivity Is the Fix

Embiricos half-jokes that AI products are hard to use because you must remember to prompt them. Users prompt maybe tens of times a day, but could benefit thousands of times — so a core Codex goal is a teammate that is helpful by default without being invoked.

  • Users prompt AI tens of times a day but could benefit thousands of times
  • If you're not prompting the model, it's probably not helping you
  • The goal is a teammate agent that is helpful by default

I like to joke today that like AI products, and it's it's a half joke, they're actually like really hard to use, because you have…

Alexander Embiricos · 14:00

probably like tens of times. But if you think of how many times people could actually get benefit from a really intelligent entity, it's thousands…

Alexander Embiricos · 14:30
#agents#product#proactivity#ux
Hot Take28:00

The Best Way for Models to Use Computers Is to Write Code

For a super assistant to do things, it needs to use a computer — and Embiricos argues the best way to do that isn't clicking or accessibility APIs, it's writing code. The implication: any agent should probably be a coding agent, even for tasks like financial analysis, and non-technical users won't even realize it.

  • Models are far more effective at doing things when they can use a computer
  • Writing code beats point-and-click or hacking accessibility APIs
  • Any agent should probably be a coding agent — code is composable and importable

it turns out the best way for models to use computers is simply to write code.

Alexander Embiricos · 29:00

maybe to the user, a non-technical user, they won't even know they're using a coding agent the same way that no one thinks about are…

Alexander Embiricos · 29:00
#agents#codex#coding#product-vision
Hot Take53:30

Ideas Still Aren't Worth Much — Bet on Deep Customer Understanding

Even as building gets cheap, Embiricos argues ideas remain overrated and execution and distribution are what's hard. If he could pick one core competency it would be a meaningful understanding of a specific customer's problems — making him bullish on vertical AI startups serving customers underserved by current tools.

  • Ideas aren't worth as much as many think; execution and distribution are still hard
  • The one competency worth having is deep understanding of a specific customer's problems
  • Founders with an underserved customer network are set; generic builders face a harder time

I still don't think ideas are worth as much as maybe some a lot of people think. I think still think execution is really hard,…

Alexander Embiricos · 53:30

if you're like good [clears throat] at building like, you know, websites, but you don't have any specific customer to build for, I think you're…

Alexander Embiricos · 55:00
#startups#distribution#vertical-ai#strategy
Hot Take1:10:30

The Underappreciated AGI Bottleneck Is Human Typing Speed

Asked about AGI timelines, Embiricos argues an underappreciated limiting factor is literally human typing and multitasking speed — how fast humans can write prompts and review the agent's output. The unlock is rebuilding systems so agents are default-useful, which he expects to start hockey-sticking early adopters' productivity next year.

  • The current limiting factor is human typing/multitasking speed on prompts and review
  • If the agent can't validate its own work, you're bottlenecked reviewing all its code
  • Rebuilding systems for default-useful agents will hockey-stick productivity — starting with early adopters next year

I think a current underappreciated limiting factor is like literally human typing speed or human multitasking speed on like writing prompts.

Alexander Embiricos · 1:11:00
#agi#agents#productivity#bottleneck

Explainer· 5

Explainer08:00

Why OpenAI Runs 'Ready, Fire, Aim' and Truly Bottoms-Up

Because OpenAI can't predict which capabilities will arrive or land, the org optimizes for humility and empirical learning over up-front planning. Embiricos contrasts this with rallying-the-ship PM work at Dropbox and his startup, noting the approach only works because of an unusually high talent bar.

  • The org is deliberately bottoms-up because capabilities and outcomes are unpredictable
  • Aim exists but is fuzzy — good conversations happen about a year out or a few weeks out, not the awkward middle
  • The model relies on a talent caliber most companies can't replicate

it's much more important for us to be very like humble and learn a lot more empirically and just try things quickly.

Alexander Embiricos · 08:30

OpenAI is like truly truly bottoms up and that's like been a learning experience for me.

Alexander Embiricos · 08:30
#product#org-design#openai#management
Explainer12:30

Codex Today Is a 'Smart Intern That Refuses to Read Slack'

Embiricos frames Codex not as autocomplete but as the beginning of a software engineering teammate. Today it's like a smart intern lacking context — so people pair with it — but the goal is an agent that participates across the whole development cycle from ideation to maintenance.

  • Codex is positioned as a software engineering teammate, not just a code writer
  • Today it lacks ambient context, so people pair with it rather than fully delegate
  • The vision spans ideation, planning, validation, deploying, and maintaining

it's a bit like this like really smart intern that like refuses to read Slack, and like doesn't check DataDog or like Sentry unless you…

Alexander Embiricos · 12:30
#codex#agents#coding#product-vision
Explainer17:30

What Unlocked Codex Growth: Pulling Back From the Cloud

The first Codex was a cloud agent with its own computer that you delegated to asynchronously — powerful but hard to adopt because of environment setup and async prompting. The unlock was meeting users where they are: an interactive IDE extension or CLI agent running in a local sandbox that turns pairing into a fast feedback loop.

  • The original cloud product was too far in the future — hard to configure and prompt
  • Async-only delegation is like a teammate you can never get on a call with
  • The unlock was landing locally in the IDE/CLI with a sandbox and an intuitive feedback loop

And we're seeing that growing, but the key unlock is actually first you need to land with users in a way that's like much more…

Alexander Embiricos · 18:30

if you hire a teammate and you ask them to do work, but they you just give them like a fresh computer from the store,…

Alexander Embiricos · 19:30
#codex#product#adoption#sandbox
Explainer22:30

An Agent Is a Stack: Model, API, and Harness — and Compaction Proves It

Embiricos reframes progress from 'just train the best model' to building the whole agent stack: the reasoning model, the API serving it, and the harness. Compaction — letting Codex run overnight or for 24 hours past its context window — required coordinated changes across all three layers.

  • An agent is a stack of model + API + harness, not just a model
  • Codex can run for very long periods, sometimes overnight or 24 hours
  • Compaction lets it exceed its context window and required all three layers to cooperate

And so, you know, for a model to work continuously for that amount of time, it's going to exceed its context window. And so, we…

Alexander Embiricos · 23:00
#agents#codex#context-window#architecture
Explainer56:30

Why the Codex Team Watches Reddit More Than Twitter

Embiricos says the Codex team may be the most user-feedback and social-media-pilled team in the space, constantly reading Reddit and Twitter and taking complaints seriously. He finds Twitter/X hypey and one-to-one, while Reddit's upvoting gives more honest, higher-signal feedback — so he increasingly watches r/codex.

  • The team constantly monitors Reddit and Twitter and takes complaints seriously
  • Twitter/X skews hypey and one-to-one; Reddit is more negative but real
  • Reddit's upvoting surfaces higher-signal feedback; he watches r/codex

Especially I think for for Twitter X, um it's a little bit more hypey. And then Reddit is a little more

Alexander Embiricos · 57:00
#codex#user-feedback#product#reddit

Story· 4

Story05:30

What Working at OpenAI Reimagined About Speed and Ambition

Embiricos, a former startup founder and Dropbox PM, says OpenAI's speed and ambition are dramatically higher than anywhere he's worked. Living through Codex's explosive scaling reset his baseline for how fast a tech product can grow, and made him far more ruthless about where he spends his time.

  • OpenAI's speed and ambition exceeded even a startup founder's expectations
  • Living through Codex's growth reset his sense of achievable pace and scale
  • The scale of impact required forces harder prioritization of his own time

By far, I would say the speed and ambition of working at OpenAI are just like dramatically more than what I can imagine.

Alexander Embiricos · 05:30
#openai#culture#startups#product
Story15:30

Codex Grew 20x Since GPT-5 and Serves Trillions of Tokens a Week

Codex has grown explosively since GPT-5's August launch — over 20x. Its models now serve many trillions of tokens a week and are OpenAI's most-served coding model, and Embiricos notes external API coding customers have started adopting the Codex models too.

  • Codex grew roughly 20x since GPT-5 launched in August
  • Codex models serve many trillions of tokens per week
  • It's the most-served coding model, now also being adopted by other API coding customers

again, the last the last thought we shared there was like we were like well over 10x since August. In fact, it's been like 20x…

Alexander Embiricos · 15:30

Also, the Codex models are serving many many trillions of tokens a week now, and it's basically like our most served coding model.

Alexander Embiricos · 16:00
#codex#growth#metrics#models
Story47:00

The Sora Android App: Built in 28 Days, #1 in the App Store

Embiricos calls the Sora Android app one of the most mind-blowing examples of acceleration. Engineers had Codex study the existing iOS app, produce plans, and implement them — reaching employees in 18 days and public GA at 28 days, with just two or three engineers, becoming the #1 app in the App Store.

  • A fully new app reached employees in 18 days and GA in 28 days total
  • Built with just two or three engineers, heavily using Codex
  • Codex excels at porting when there's an existing app (iOS) to reference; the result hit #1 in the App Store

the Sora Android app, right? Like a fully new app, we built it in 18 days. It went from like zero to launch to employees.…

Alexander Embiricos · 47:30

imagine there's the number one app in the App Store with like a handful of engineers. Uh I think it was like two or three…

Alexander Embiricos · 48:30
#codex#sora#acceleration#case-study
Story49:00

Atlas: From 2–3 Engineers for 2–3 Weeks to One Engineer, One Week

Building a browser is hard, but the Atlas team became Codex power users. Embiricos relays that work that previously would have taken two to three engineers two to three weeks now takes one engineer one week — and the team is now using Codex to make the model better on Windows and PowerShell.

  • Atlas engineers use Codex for 'absolutely everything'
  • Roughly a 6x acceleration: 2–3 engineers × 2–3 weeks down to 1 engineer × 1 week
  • Last week's model was the first to natively understand PowerShell for the Windows version

before this would have taken us like 2 to 3 weeks for two to three engineers. And now it's like one engineer, one week.

Alexander Embiricos · 49:30

And it's basically like, we use Codex for absolutely everything.

Alexander Embiricos · 51:00
#codex#atlas#acceleration#case-study

Takeaway· 2

Takeaway34:30

Writing Code Is the Fun Part — Reviewing AI Code Isn't

Embiricos observes that writing code is one of the most fun parts of the job, but agents shift engineers into reviewing AI-written code, which is far less fun. The product response: ship code-review features and design choices (like showing the image preview before the diff) that keep humans empowered.

  • Coding agents shift engineers from writing to reviewing — the less fun part
  • The product fix is helping people gain confidence in AI-written code
  • Micro-decisions matter: show the image preview before the diff to keep humans in flow

it turns out writing code is actually one of the most fun parts of software engineering for many software engineers. And so then you end…

Alexander Embiricos · 34:30
#codex#code-review#developer-experience#product
Takeaway1:02:30

The Best Way to Try Codex: Give It Your Hardest Task

Unlike other coding agents where you might start with something trivial, Embiricos says the best way to try Codex is to hand it your hardest task — a real, gnarly bug in your enormous, imperfect codebase. It's built as a professional tool for your hardest single problem, not a demo toy.

  • Give Codex a real, hard task rather than dumbing it down to something trivial
  • It's built as a professional tool for high-quality work in large, imperfect codebases
  • Start with one hard question or task (e.g. a mystery bug), not an entire business

the best way to try Codex is to give it your hardest tasks, which is a little different than some of the other coding agents.

Alexander Embiricos · 1:03:00
#codex#tips#adoption#coding