LLenny's Podcast
← All episodes
Alexander Embiricos (OpenAI Codex Product Lead)14 December 2025

Why humans are AI’s biggest bottleneck (and what’s coming in 2026)

6Frameworks
15Insights

Episode overview

OpenAI Codex product lead Alexander Embiricos explains how Codex evolved from a cloud-only asynchronous agent into an interactive coding teammate embedded in developers’ existing workflows. He attributes its rapid growth and performance to jointly developing the reasoning model, API, harness, sandbox, and compaction capabilities, while measuring early retention and closely monitoring user feedback. The conversation explores how Codex accelerated projects such as the Sora Android app and Atlas, why code review and validation have replaced code generation as important bottlenecks, and how proactive coding agents could become general-purpose assistants. Embiricos argues that human prompting, multitasking, and review speed currently constrain AI-driven productivity, with early adopters likely to see sharp gains before complex established organizations do.

Key ideas

  • Codex is intended to evolve from an interactive coding tool into a proactive software-engineering teammate that can participate across planning, implementation, validation, deployment, and maintenance.
  • Codex's adoption accelerated when OpenAI supplemented its cloud delegation model with local IDE and CLI experiences that fit existing developer workflows and provide immediate feedback.
  • OpenAI develops Codex's model, API, harness, sandbox, and compaction system together, enabling long-running tasks and optimization for a shell-based agent workflow.
  • Coding may become a core competency of general-purpose agents because writing code is a reliable, composable way for models to operate computers and reuse learned procedures.
  • As agents generate more code, validation and code review become major bottlenecks; products must help agents verify their own work while keeping humans meaningfully in control.
  • Internal Codex use reportedly helped OpenAI build the Sora Android app in 18 days for employees and launch publicly 10 days later, while also substantially accelerating Atlas engineering.
  • Cheaper code generation raises the relative importance of customer understanding, distribution, coherent execution, systems engineering, and effective team communication.
  • Embiricos expects productivity gains to become hockey-stick-like when agents are useful by default and can validate their work without constant human prompting and manual review.

Transcript available · source text is retained privately and is not published

Frameworks in this episode

People & resources mentioned

Attributed to the moment in the episode. Timestamps are approximate.

People · 12

  • Nick TurleyMentionshead of ChachiPT, Nick Turley was like almost going to become professional jazz pianist.

    In the words of Nick Turley, head of ChatGPT and former podcast guest

  • Alexander EmbiricosMentionsToday my guest is Alexander Imbiricos, product lead for Codex, OpenAI's incredibly popular and powerful coding agent.

    Today my guest is Alexander Imbiricos, product lead for Codex, OpenAI's incredibly popular and powerful coding agent.

  • Lenny RachitskyMentionstoday we've got another very special compilation episode something I've been pulling on more and more with the podcast and the newsletter

    Today my guest is Alexander Imbiricos, product lead for Codex, OpenAI's incredibly popular and powerful coding agent.

  • Kevin WeilMentionsWe got connected through Kevin Wheel, who is former CPO at OpenAI, now head of science at OpenAI.

    Similarly, Kevin Weil, OpenAI CPO, said Alex is simply the best.

  • Dennis YangMentionsI have this good friend uh his name's Dennis Yang he works at chime

    A huge thank you to Ed Bay, Nick Turley, and Dennis Yang for suggesting topics for this conversation.

  • Ed BayMentionsmentioned

    A huge thank you to Ed Bay, Nick Turley, and Dennis Yang for suggesting topics for this conversation.

  • Andrej KarpathyMentionsI remember Karpathy tweeted that he just like has never seen a model like this.

    I remember Karpathy tweeted that he just like has never seen a model like this.

  • Michael TruellMentionsMichael is co-founder and CEO of Any Sphere, the company behind Cursor.

    When I had Michael Turalda, CEO of Cursor on the podcast

  • Scott BelskyCoinedHe's a former founder starting a company called Behance that he sold to Adobe where he worked up the ranks to chief product officer

    Scott Belsky talks about this idea of like compressing the talent stack.

  • Iain M. BanksCoinedI think it's Ian Banks is the name of the author.

    I think it's Ian Banks is the name of the author.

  • Andreas EmbiricosMentionsThe influential Greek poet and psychoanalyst Andreas Imbrios.

    The influential Greek poet and psychoanalyst Andreas Imbrios.

  • George EmbiricosMentionsThe wealthy shipping magnate and art collector George Imbrios.

    The wealthy shipping magnate and art collector George Imbrios.

Resources · 39

  • Lenny's NewsletterCoinednewsletter · Lenny Rachitsky

    if you become an annual subscriber of my newsletter

  • DropboxMentionscompany · Drew Houston and Arash Ferdowsi

    Before that, you were a product manager at Dropbox.

  • OpenAIMentionscompany

    you joined OpenAI about a year ago.

  • Visual Studio CodeMentionssoftware · Microsoft

    it's an IDE extension, like a VS Code extension

  • CodexCoinedsoftware · OpenAI

    Codex is Open AI's coding agent.

  • DatadogMentionscompany · Datadog

    doesn't check DataDog or like Sentry unless you ask it to.

  • SentryMentionssoftware · Sentry

    doesn't check DataDog or like Sentry unless you ask it to.

  • CursorMentionssoftware · Anysphere

    When people think about Cursor and even Cloud Code, it's like IDE that helps you code

  • GPT-5Mentionssoftware · OpenAI

    Codex has been growing like absolutely explosively since the launch of GPT-5 back in August.

  • Claude CodeMentionssoftware · Anthropic

    It just felt like Cloud Code was killing it.

  • Codex CloudCoinedsoftware · OpenAI

    we launched our first version of Codex, uh which was Codex Cloud.

  • GPT-5.1 Codex MaxCoinedsoftware · OpenAI

    just last Wednesday, we shipped GPT-5.1 Codex Max.

  • ChatGPTMentionssoftware · OpenAI

    you have chat, ChatGPT, and that is a tool that's like ubiquitously available to like everyone.

  • GitHub CopilotMentionssoftware · GitHub

    the first time we used in the the brand Codex at OpenAI was actually the model powering GitHub Copilot.

  • ChatGPT AtlasCoinedsoftware · OpenAI

    when I think about launching a browser, which we did with Atlas

  • GooseMentionssoftware · Block

    they they have this product called Goose, which is their own internal agent thing.

  • SlackUsessoftware · Slack Technologies

    we have a Slack integration for Codex.

  • SoraMentionssoftware · OpenAI

    we recently shipped the Sora Android app.

  • Sora Android appUsesproduct · OpenAI

    the Sora Android app, right? Like a fully new app, we built it in 18 days.

  • App StoreMentionssoftware · Apple

    it became the number one app in the App Store.

  • PowerShellMentionssoftware · Microsoft

    the model we we shipped last week is the first model that natively understands PowerShell.

  • XUseswebsite · X Corp.

    a few of us are like constantly on Reddit and Twitter.

  • RedditUseswebsite · Steve Huffman, Alexis Ohanian, and Aaron Swartz

    a few of us are like constantly on Reddit and Twitter.

  • r/CodexUseswebsite

    like r/Codex is is there.

  • HaloMentionsother · Bungie

    I don't know if you've played like, I don't know, say Halo, right?

  • CodexRecommendssoftware · OpenAI

    the best way to try Codex is to give it your hardest tasks

  • CodexUsessoftware · OpenAI

    Codex writes a lot of the code that helps like manage its training runs, the key infrastructure.

  • SAPMentionscompany · SAP

    Let's say you work in SAP. Like they have many like complex systems

  • The CultureRecommendsbook · Iain M. Banks

    I'm sure this has been recommended before, but the Culture.

  • The Lord of the RingsMentionsbook · J. R. R. Tolkien

    I know you're reading, you mentioned before we started recording, Lord of the Rings right now.

  • A Fire Upon the DeepRecommendsbook · Vernor Vinge

    have you read Fire Upon the Deep?

  • Neon Genesis EvangelionMentionstv show · Hideaki Anno

    if you look at like some older anime like that started the genre like, you know, there was there was like Evangelion

  • Jujutsu KaisenRecommendstv show · Gege Akutami

    there's an anime called Jujutsu Kaisen, which I really like.

  • AkiraMentionsfilm · Katsuhiro Otomo

    there was there was like Evangelion uh or Akira.

  • TeslaUsescompany · Tesla, Inc.

    And then recently We got a Tesla instead.

  • TeslaRecommendscompany · Tesla, Inc.

    I have to say that I find the Tesla software like quite inspiring.

  • Radical CandorMentionsbook · Kim Scott

    Radical candor. Oh yeah, yeah, right.

  • AndrosMentionsplace

    we all came from this island called Andros.

  • Lenny's Podcast websiteMentionswebsite · Lenny Rachitsky

    You can find all past episodes or learn more about the show at lennyspodcast.com.

Spot an error or want something removed? Request a correction or removal.

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Hot Take· 4

Hot Take14:00

AI Products Are Actually Hard to Use — And Proactivity Is the Fix

Embiricos half-jokes that AI products are hard to use because you must remember to prompt them. Users prompt maybe tens of times a day, but could benefit thousands of times — so a core Codex goal is a teammate that is helpful by default without being invoked.

  • Users prompt AI tens of times a day but could benefit thousands of times
  • If you're not prompting the model, it's probably not helping you
  • The goal is a teammate agent that is helpful by default

I like to joke today that like AI products, and it's it's a half joke, they're actually like really hard to use, because you have…

Alexander Embiricos · 14:00

probably like tens of times. But if you think of how many times people could actually get benefit from a really intelligent entity, it's thousands…

Alexander Embiricos · 14:30
#agents#product#proactivity#ux
Hot Take28:00

The Best Way for Models to Use Computers Is to Write Code

For a super assistant to do things, it needs to use a computer — and Embiricos argues the best way to do that isn't clicking or accessibility APIs, it's writing code. The implication: any agent should probably be a coding agent, even for tasks like financial analysis, and non-technical users won't even realize it.

  • Models are far more effective at doing things when they can use a computer
  • Writing code beats point-and-click or hacking accessibility APIs
  • Any agent should probably be a coding agent — code is composable and importable

it turns out the best way for models to use computers is simply to write code.

Alexander Embiricos · 29:00

maybe to the user, a non-technical user, they won't even know they're using a coding agent the same way that no one thinks about are…

Alexander Embiricos · 29:00
#agents#codex#coding#product-vision
Hot Take53:30

Ideas Still Aren't Worth Much — Bet on Deep Customer Understanding

Even as building gets cheap, Embiricos argues ideas remain overrated and execution and distribution are what's hard. If he could pick one core competency it would be a meaningful understanding of a specific customer's problems — making him bullish on vertical AI startups serving customers underserved by current tools.

  • Ideas aren't worth as much as many think; execution and distribution are still hard
  • The one competency worth having is deep understanding of a specific customer's problems
  • Founders with an underserved customer network are set; generic builders face a harder time

I still don't think ideas are worth as much as maybe some a lot of people think. I think still think execution is really hard,…

Alexander Embiricos · 53:30

if you're like good [clears throat] at building like, you know, websites, but you don't have any specific customer to build for, I think you're…

Alexander Embiricos · 55:00
#startups#distribution#vertical-ai#strategy
Hot Take1:10:30

The Underappreciated AGI Bottleneck Is Human Typing Speed

Asked about AGI timelines, Embiricos argues an underappreciated limiting factor is literally human typing and multitasking speed — how fast humans can write prompts and review the agent's output. The unlock is rebuilding systems so agents are default-useful, which he expects to start hockey-sticking early adopters' productivity next year.

  • The current limiting factor is human typing/multitasking speed on prompts and review
  • If the agent can't validate its own work, you're bottlenecked reviewing all its code
  • Rebuilding systems for default-useful agents will hockey-stick productivity — starting with early adopters next year

I think a current underappreciated limiting factor is like literally human typing speed or human multitasking speed on like writing prompts.

Alexander Embiricos · 1:11:00
#agi#agents#productivity#bottleneck

Explainer· 5

Explainer08:00

Why OpenAI Runs 'Ready, Fire, Aim' and Truly Bottoms-Up

Because OpenAI can't predict which capabilities will arrive or land, the org optimizes for humility and empirical learning over up-front planning. Embiricos contrasts this with rallying-the-ship PM work at Dropbox and his startup, noting the approach only works because of an unusually high talent bar.

  • The org is deliberately bottoms-up because capabilities and outcomes are unpredictable
  • Aim exists but is fuzzy — good conversations happen about a year out or a few weeks out, not the awkward middle
  • The model relies on a talent caliber most companies can't replicate

it's much more important for us to be very like humble and learn a lot more empirically and just try things quickly.

Alexander Embiricos · 08:30

OpenAI is like truly truly bottoms up and that's like been a learning experience for me.

Alexander Embiricos · 08:30
#product#org-design#openai#management
Explainer12:30

Codex Today Is a 'Smart Intern That Refuses to Read Slack'

Embiricos frames Codex not as autocomplete but as the beginning of a software engineering teammate. Today it's like a smart intern lacking context — so people pair with it — but the goal is an agent that participates across the whole development cycle from ideation to maintenance.

  • Codex is positioned as a software engineering teammate, not just a code writer
  • Today it lacks ambient context, so people pair with it rather than fully delegate
  • The vision spans ideation, planning, validation, deploying, and maintaining

it's a bit like this like really smart intern that like refuses to read Slack, and like doesn't check DataDog or like Sentry unless you…

Alexander Embiricos · 12:30
#codex#agents#coding#product-vision
Explainer17:30

What Unlocked Codex Growth: Pulling Back From the Cloud

The first Codex was a cloud agent with its own computer that you delegated to asynchronously — powerful but hard to adopt because of environment setup and async prompting. The unlock was meeting users where they are: an interactive IDE extension or CLI agent running in a local sandbox that turns pairing into a fast feedback loop.

  • The original cloud product was too far in the future — hard to configure and prompt
  • Async-only delegation is like a teammate you can never get on a call with
  • The unlock was landing locally in the IDE/CLI with a sandbox and an intuitive feedback loop

And we're seeing that growing, but the key unlock is actually first you need to land with users in a way that's like much more…

Alexander Embiricos · 18:30

if you hire a teammate and you ask them to do work, but they you just give them like a fresh computer from the store,…

Alexander Embiricos · 19:30
#codex#product#adoption#sandbox
Explainer22:30

An Agent Is a Stack: Model, API, and Harness — and Compaction Proves It

Embiricos reframes progress from 'just train the best model' to building the whole agent stack: the reasoning model, the API serving it, and the harness. Compaction — letting Codex run overnight or for 24 hours past its context window — required coordinated changes across all three layers.

  • An agent is a stack of model + API + harness, not just a model
  • Codex can run for very long periods, sometimes overnight or 24 hours
  • Compaction lets it exceed its context window and required all three layers to cooperate

And so, you know, for a model to work continuously for that amount of time, it's going to exceed its context window. And so, we…

Alexander Embiricos · 23:00
#agents#codex#context-window#architecture
Explainer56:30

Why the Codex Team Watches Reddit More Than Twitter

Embiricos says the Codex team may be the most user-feedback and social-media-pilled team in the space, constantly reading Reddit and Twitter and taking complaints seriously. He finds Twitter/X hypey and one-to-one, while Reddit's upvoting gives more honest, higher-signal feedback — so he increasingly watches r/codex.

  • The team constantly monitors Reddit and Twitter and takes complaints seriously
  • Twitter/X skews hypey and one-to-one; Reddit is more negative but real
  • Reddit's upvoting surfaces higher-signal feedback; he watches r/codex

Especially I think for for Twitter X, um it's a little bit more hypey. And then Reddit is a little more

Alexander Embiricos · 57:00
#codex#user-feedback#product#reddit

Story· 4

Story05:30

What Working at OpenAI Reimagined About Speed and Ambition

Embiricos, a former startup founder and Dropbox PM, says OpenAI's speed and ambition are dramatically higher than anywhere he's worked. Living through Codex's explosive scaling reset his baseline for how fast a tech product can grow, and made him far more ruthless about where he spends his time.

  • OpenAI's speed and ambition exceeded even a startup founder's expectations
  • Living through Codex's growth reset his sense of achievable pace and scale
  • The scale of impact required forces harder prioritization of his own time

By far, I would say the speed and ambition of working at OpenAI are just like dramatically more than what I can imagine.

Alexander Embiricos · 05:30
#openai#culture#startups#product
Story15:30

Codex Grew 20x Since GPT-5 and Serves Trillions of Tokens a Week

Codex has grown explosively since GPT-5's August launch — over 20x. Its models now serve many trillions of tokens a week and are OpenAI's most-served coding model, and Embiricos notes external API coding customers have started adopting the Codex models too.

  • Codex grew roughly 20x since GPT-5 launched in August
  • Codex models serve many trillions of tokens per week
  • It's the most-served coding model, now also being adopted by other API coding customers

again, the last the last thought we shared there was like we were like well over 10x since August. In fact, it's been like 20x…

Alexander Embiricos · 15:30

Also, the Codex models are serving many many trillions of tokens a week now, and it's basically like our most served coding model.

Alexander Embiricos · 16:00
#codex#growth#metrics#models
Story47:00

The Sora Android App: Built in 28 Days, #1 in the App Store

Embiricos calls the Sora Android app one of the most mind-blowing examples of acceleration. Engineers had Codex study the existing iOS app, produce plans, and implement them — reaching employees in 18 days and public GA at 28 days, with just two or three engineers, becoming the #1 app in the App Store.

  • A fully new app reached employees in 18 days and GA in 28 days total
  • Built with just two or three engineers, heavily using Codex
  • Codex excels at porting when there's an existing app (iOS) to reference; the result hit #1 in the App Store

the Sora Android app, right? Like a fully new app, we built it in 18 days. It went from like zero to launch to employees.…

Alexander Embiricos · 47:30

imagine there's the number one app in the App Store with like a handful of engineers. Uh I think it was like two or three…

Alexander Embiricos · 48:30
#codex#sora#acceleration#case-study
Story49:00

Atlas: From 2–3 Engineers for 2–3 Weeks to One Engineer, One Week

Building a browser is hard, but the Atlas team became Codex power users. Embiricos relays that work that previously would have taken two to three engineers two to three weeks now takes one engineer one week — and the team is now using Codex to make the model better on Windows and PowerShell.

  • Atlas engineers use Codex for 'absolutely everything'
  • Roughly a 6x acceleration: 2–3 engineers × 2–3 weeks down to 1 engineer × 1 week
  • Last week's model was the first to natively understand PowerShell for the Windows version

before this would have taken us like 2 to 3 weeks for two to three engineers. And now it's like one engineer, one week.

Alexander Embiricos · 49:30

And it's basically like, we use Codex for absolutely everything.

Alexander Embiricos · 51:00
#codex#atlas#acceleration#case-study

Takeaway· 2

Takeaway34:30

Writing Code Is the Fun Part — Reviewing AI Code Isn't

Embiricos observes that writing code is one of the most fun parts of the job, but agents shift engineers into reviewing AI-written code, which is far less fun. The product response: ship code-review features and design choices (like showing the image preview before the diff) that keep humans empowered.

  • Coding agents shift engineers from writing to reviewing — the less fun part
  • The product fix is helping people gain confidence in AI-written code
  • Micro-decisions matter: show the image preview before the diff to keep humans in flow

it turns out writing code is actually one of the most fun parts of software engineering for many software engineers. And so then you end…

Alexander Embiricos · 34:30
#codex#code-review#developer-experience#product
Takeaway1:02:30

The Best Way to Try Codex: Give It Your Hardest Task

Unlike other coding agents where you might start with something trivial, Embiricos says the best way to try Codex is to hand it your hardest task — a real, gnarly bug in your enormous, imperfect codebase. It's built as a professional tool for your hardest single problem, not a demo toy.

  • Give Codex a real, hard task rather than dumbing it down to something trivial
  • It's built as a professional tool for high-quality work in large, imperfect codebases
  • Start with one hard question or task (e.g. a mystery bug), not an entire business

the best way to try Codex is to give it your hardest tasks, which is a little different than some of the other coding agents.

Alexander Embiricos · 1:03:00
#codex#tips#adoption#coding