LLenny's Podcast
← All episodes
11 January 2026

Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google, and Amazon

4Frameworks
15Insights

Episode overview

Aishwarya Reganti and Kiriti Badam explain why AI products require a different development approach: both user inputs and model outputs are nondeterministic, while increasing an agent's agency reduces human control. Drawing on more than 50 deployments, they advocate starting with narrow, low-risk capabilities and using continuous calibration, evaluation, production monitoring, and human feedback to expand autonomy safely. They also argue that successful adoption depends on hands-on leadership, an empowering culture, deep workflow understanding, and persistence through repeated experimentation.

Key ideas

  • AI products differ from conventional software because natural-language inputs and probabilistic model outputs make behavior difficult to predict.
  • Teams should begin with low agency and high human control, then increase autonomy only as system behavior becomes understood and trusted.
  • A problem-first approach prevents teams from becoming distracted by sophisticated agents and losing sight of the customer workflow they need to improve.
  • Evaluations and production monitoring serve complementary purposes: evals test known failure modes, while live signals reveal unexpected behavior and new error patterns.
  • The Continuous Calibration, Continuous Development framework combines scoped capabilities, curated data, evaluation metrics, deployment, behavior analysis, and iterative fixes.
  • Successful enterprise AI adoption requires hands-on leaders, a culture that empowers subject-matter experts, and careful decomposition of workflows across AI, deterministic code, and humans.
  • Durable advantage comes from building feedback flywheels and accumulating hard-won knowledge through repeated experimentation—summarized as “pain is the new moat.”
  • The guests expect proactive background agents, coding agents, and richer multimodal systems to create substantial value as they gain better access to work context.

Transcript available · source text is retained privately and is not published

Frameworks in this episode

People & resources mentioned

Attributed to the moment in the episode. Timestamps are approximate.

People · 14

  • Aishwarya RegantiMentionsAsh was an early AI researcher at Alexa and Microsoft and has published over 35 research papers.

    Today my guests are Aishwaria Raanti and Kiti Bottom.

  • Kiriti BadamMentionsKiti works on codecs at OpenAI and has spent the last decade building AI and ML infrastructure at Google and at Kumo.

    Today my guests are Aishwaria Raanti and Kiti Bottom.

  • Lenny RachitskyMentionstoday we've got another very special compilation episode something I've been pulling on more and more with the podcast and the newsletter

    Today my guests are Aishwaria Raanti and Kiti Bottom.

  • Matei ZahariaMentionsmentioned

    I recently read this paper from a bunch of folks at UC Berkeley. um basically mate Zahara Stoker and the folks at data bricks

  • Gajen KandiahMentionsIdentified as the CEO of Rackspace and a former colleague of Aishwarya Reganti.

    I used to work with the CEO of now Rackspace Gajen.

  • Dan ShipperMentionsDan is the co-founder and CEO of Every which is a company that is at the very bleeding edge of what is possible with AI.

    I had Dan Shipper on the podcast and they work with a bunch of companies helping them adopt AI

  • Martin FowlerCoinedmentioned

    Martin Fowler at some point had this term called semantic diffusion back in the 2000s.

  • Demis HassabisMentionsReferenced for an interview about testing AI against historical scientific breakthroughs.

    I just saw Demis from Deep Mind AI, Google, whatever they call the whole or uh talking about this

  • Jason LemkinMentionsJason created and runs SaaStr and previously co-founded and led EchoSign before its acquisition by Adobe.

    I just had Jason Lumpkit on the podcast. He's um uh very smart on sales, go to market, run Zaster

  • Paul KalanithiCoinedDescribed as a neurosurgeon diagnosed with lung cancer in his early thirties.

    For me, it's this book called When Breath Becomes Air, Lenny. It was written by Paul Kalaniti.

  • SocratesCoinedmentioned

    he was kind of arguing against a very popular quote by Socrates which is the unexamined life is not worth living

  • Vernor VingeCoinedmentioned

    you got to read this book called A Fire Upon the Deep by uh, Vernon Vege.

  • Noah SmithRecommendsWriter of the Noahpinion newsletter

    I saw Noah Smith on his newsletter recommend this book

  • Steve JobsCoinedSteve Jobs is just uh the bar he held for the company and for technical talent and for excellence was not uh wavering.

    I really like this quote from Steve Jobs that you can only connect the dots looking backwards

Resources · 32

  • GoogleMentionscompany · Google

    Kiti works on codecs at OpenAI and has spent the last decade building AI and ML infrastructure at Google and at Kumo.

  • AlexaMentionssoftware · Amazon

    Ash was an early AI researcher at Alexa and Microsoft and has published over 35 research papers.

  • MicrosoftMentionscompany · Bill Gates and Paul Allen

    Ash was an early AI researcher at Alexa and Microsoft and has published over 35 research papers.

  • AmazonMentionscompany · Jeff Bezos

    Together, they've led and supported over 50 AI product deployments across companies like Amazon, Data Bricks, OpenAI, Google

  • DatabricksMentionscompany · Ali Ghodsi, Matei Zaharia and co-founders

    Together, they've led and supported over 50 AI product deployments across companies like Amazon, Data Bricks, OpenAI, Google

  • MavenUseswebsite · Maven

    Together, they also teach the number one rated AI course on Maven

  • Booking.comMentionscompany

    Think of something like booking.com right you um you have an intention that uh you want to make a booking in San Francisco

  • OpenAIUsescompany

    OpenAF faced the exact same thing when we were launching products and there was like a huge spike of uh support volume

  • RackspaceMentionscompany

    I used to work with the CEO of now Rackspace Gajen.

  • ChatGPTMentionssoftware · OpenAI

    For example, in chat GPD, right? Like if you are uh liking the answer you can actually give a thumbs up

  • Artificial AnalysisUseswebsite

    No, we just checked LM arena and artificial analysis.

  • LM ArenaUseswebsite

    No, we just checked LM arena and artificial analysis.

  • CodexUsessoftware · OpenAI

    So CEX we have like this balanced approach of like you know you need to have eval and you need to definitely listen

  • Air CanadaMentionscompany

    I think Air Canada had this thing where um one of their agents predicted or hallucinated a policy

  • Continuous Calibration, Continuous DevelopmentCoinedother · Aishwarya Reganti and Kiriti Badam

    that's why we came up with this idea of continuous calibration continuous development.

  • GPT-4oMentionssoftware · OpenAI

    GPD 40 doesn't exist anymore or it's going to be deprecated in APIs as well.

  • GPT-5Mentionssoftware · OpenAI

    So most companies that were using 40 should switch to five and five has very different properties.

  • ChatGPT PulseMentionsproduct · OpenAI

    we already do this in terms of charge GPT pulse which kind of gives you this daily update of things you might care about

  • LinearMentionssoftware · Linear

    a coding agent which says that like okay I have fixed five of your linear tickets and here are the patches

  • DeepMindMentionscompany

    I just saw Demis from Deep Mind AI, Google, whatever they call the whole or uh talking about this

  • GenieMentionssoftware · Google DeepMind

    combining the image model work, the LLM, and also their world model stuff, Genie, I think is what it's called.

  • SalesforceUsescompany · Salesforce

    one of the agents, it was just tracking everyone's updates to Salesforce and kind of uh updating it automatically for them based on their calls.

  • When Breath Becomes AirRecommendsbook · Paul Kalanithi

    For me, it's this book called When Breath Becomes Air, Lenny. It was written by Paul Kalaniti.

  • Remembrance of Earth's PastRecommendsbook · Liu Cixin

    I uh really like this three body problem series. Uh it's like a three book series.

  • A Fire Upon the DeepRecommendsbook · Vernor Vinge

    you got to read this book called A Fire Upon the Deep by uh, Vernon Vege.

  • Silicon ValleyRecommendsplace · Mike Judge, John Altschuler, and Dave Krinsky

    I started re-watching Silicon Valley, and I think it's so true. It's so timeless.

  • Clair Obscur: Expedition 33Recommendsproduct · Sandfall Interactive

    there's this game I picked up recently called Expedition 33.

  • Wispr FlowRecommendssoftware · Wispr AI

    For me, it's Whisper Flow. I think I've been using it quite a bit and I didn't know I needed it so much.

  • RaycastRecommendssoftware · Raycast

    Uh so I feel like a recast has been amazing.

  • caffeinateRecommendstool · Apple

    And caffeinate is another thing that I've recently discovered from my teammates. It helps you like prevent Mac from sleeping.

  • LinkedInUseswebsite · LinkedIn

    I write a lot on LinkedIn.

  • Building Enterprise AI ProductsCoinedcourse · Aishwarya Reganti and Kiriti Badam

    we also run a super popular course. We'll leave a link to it on building enterprise AI products.

Spot an error or want something removed? Request a correction or removal.

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster38:00

"Evals" Now Means Five Different Things (Semantic Diffusion)

The guest argues the word "evals" has been butchered into meaning error analysis, expert notes, LLM judges, feedback loops, and even public benchmarks. Invoking Martin Fowler's term "semantic diffusion," they warn that checking LM Arena is not doing evals. The real, agreed-upon goal is building an actionable feedback loop.

  • Data labelers, PMs, and researchers all mean different things by "evals"
  • Martin Fowler's "semantic diffusion": a term gets diluted until its meaning is lost
  • Reading LM Arena or benchmarks is not the same as building your own eval dataset
  • Whether you need an LLM judge depends entirely on context; don't over-prescribe

someone comes up with a term everybody starts butchering it with their own definitions and then you kind of lose the actual definition of it.

39:30

You're not doing eval. That's not eval. Those are model.

41:00
#evals#terminology#llm-judge#benchmarks

Hot Take· 4

Hot Take31:00

"One-Click Agents Are Pure Marketing"

The guest is openly skeptical of vendors promising a one-click agent that delivers gains in days. Enterprise data and infrastructure are too messy for that, and the real work is building a flywheel that improves over time. Meaningful ROI, even with the best data layer, takes four to six months.

  • Vendors promising instant results are ignoring messy enterprise data and tech debt
  • Prefer a company that builds a learning pipeline over one selling out-of-the-box replacement
  • Significant ROI typically takes four to six months even with good infrastructure
  • The edge is the flywheel you build, not being first to deploy an agent

if someone's selling you one click agents, it's it's pure marketing. You don't want to buy into that.

31:30
#agents#enterprise#vendor-hype#flywheel
Hot Take1:01:30

Multi-Agent Systems Are Misunderstood, Not Overhyped

The guest says multi-agent systems are misunderstood rather than overrated. A supervisor agent delegating to sub-agents is a proven, successful pattern. What doesn't work is splitting responsibilities by function and expecting the agents to coordinate on their own in a peer-to-peer "gossip protocol" — current models aren't there yet.

  • Supervisor-with-sub-agents is a genuinely successful pattern
  • Peer-to-peer, self-coordinating agents dividing work by function is the misunderstood part
  • Current model capabilities aren't ready for gossip-protocol multi-agent apps
  • In customer support especially, letting agents talk peer-to-peer makes guardrails unmanageable

if you're building a supervisor agent and there are like sub agents that actually do the work for the super agent supervisor agent that is…

1:02:00

coming with this notion of I'm going to divide the responsibilities based on functionality and somehow uh expect all of that to work together in…

1:02:00
#multi-agent#architecture#agents
Hot Take1:02:30

Coding Agents Are Still Underrated

Despite all the Twitter and Reddit chatter, the guest insists coding agents are underrated. Talk to engineers at ordinary companies, especially outside the Bay Area, and the real-world impact is huge while penetration remains very low. They expect 2025-2026 to create enormous value optimizing these processes.

  • Online hype overstates how far coding-agent adoption has actually spread
  • Outside the Bay Area, real-world penetration is very low despite high potential impact
  • 2025-2026 is expected to be a big value-creation window for coding agents

I still feel coding agents are underrated

1:02:30

talking to an engineer in like any random company uh especially outside of Bay Area you you can see like the amount of impact this…

1:03:00
#coding-agents#adoption#predictions
Hot Take1:12:00

"Pain Is the New Moat"

The guest's favorite concept: in a field with no textbook, the moat isn't being first or having a flashy feature — it's the hard-won knowledge from iterating through pain. Successful companies figured out their non-negotiables and traded them off against model capabilities, and that lived experience is what competitors can't copy.

  • There's no known path or textbook, so knowledge comes from painful iteration
  • Winners identified their non-negotiables and traded them against model capabilities
  • The accumulated org-wide lived experience is the real defensibility
  • Persistence matters because information is now at your fingertips to learn anything

I I like to call it like pain is the new mode

1:12:30

They went through the pain of understanding what are the set of non-negotiable things

1:13:00
#moat#strategy#persistence

Explainer· 3

Explainer07:30

The Two Ways Building AI Products Is Fundamentally Different

The guests argue AI products break from traditional software in two ways most people ignore. First, non-determinism: both the input (users express intent in unlimited natural-language ways) and the output (a probabilistic, black-box LLM) are unpredictable. Second, the agency-control trade-off: giving an agent more autonomy means relinquishing control, so the agent must earn trust before you let it decide.

  • Traditional software has a well-mapped decision engine; AI replaces that layer with a fluid natural-language interface
  • You don't know how the user will behave (input) or how the LLM will respond (output)
  • Every increase in agency is a decrease in your control
  • An agent should build up trust over time before being given more decision-making power

one of them that most people tend to ignore is the non-determinism

08:00

every time you hand over decision-m capabilities or autonomy to agentic systems, you're kind of relinquishing some amount of control on your end

09:30
#ai-products#non-determinism#agents#product-design
Explainer25:30

The Success Triangle: Every Tech Problem Is a People Problem First

The guest frames successful AI companies along three dimensions rather than pure technology: great leaders, good culture, and technical progress. Leaders must relearn intuitions built over a decade, and culture must move from FOMO and replacement fear toward empowerment so subject-matter experts actually cooperate.

  • Three dimensions of success: great leaders, good culture, technical progress
  • Leaders' 10-15 year intuitions must be relearned and they must be vulnerable enough to do it
  • AI adoption almost always needs a top-down push; bottom-up rarely gets leader buy-in
  • A culture of empowerment (not replacement fear) keeps subject-matter experts engaged

Every technology problem is a people problem first.

25:30

great leaders, good culture and technical progress

26:00
#leadership#culture#ai-transformation
Explainer33:30

Evals vs Production Monitoring Is a False Dichotomy

The guest rejects the idea that either evals or production monitoring alone will solve everything. Evals encode your product knowledge into datasets of what should and shouldn't happen; production monitoring surfaces real user behavior through explicit and implicit signals. You need both because each catches errors the other can't.

  • Evals are your product thinking turned into datasets of expected right/wrong behavior
  • Production monitoring reads real usage via explicit and implicit signals
  • Regenerating an answer (not just thumbs-down) is an implicit signal the output missed
  • Evals only catch errors you already anticipated; monitoring finds emerging failure patterns

there's this false dichotomy of like there's either eval is going to solve everything or online monitoring or production monitoring is going to solve everything

33:30
#evals#monitoring#reliability#feedback-loops

Story· 3

Story26:30

The CEO Who Blocks 4-6am Every Day to Catch Up on AI

The guest describes working with the CEO of what is now Rackspace, who blocks 4-6am daily with no meetings just to consume the latest AI content, plus weekend vibe-coding sessions. The point: leaders must get hands-on to rebuild their intuitions and accept they may be the dumbest person in the room.

  • The CEO reserved 4-6am daily, meeting-free, to keep up with AI
  • He curated two or three trusted sources and bounced questions off AI experts
  • Hands-on time is about rebuilding intuition, not implementing things personally

have this block every day in the morning which would say catching up with AI 4 to 6:00 a.m.

26:30

you probably are the dumbest person in the room and you want to learn from everyone

27:00
#leadership#habits#learning
Story47:00

When Air Canada's Agent Invented a Refund Policy

As a cautionary example of over-autonomous agents, the guest cites Air Canada, whose agent hallucinated a refund policy that wasn't in the original playbook, and the company was legally forced to honor it. Alongside their own customer-support product they had to shut down for endless hot-fixing, it motivates building so you don't lose customer trust or let an AI make dangerous decisions.

  • Air Canada's agent hallucinated a refund policy outside its playbook and legal forced compliance
  • The guests once shut down a customer-support product because emerging problems needed endless hot-fixes
  • Big end-to-end agents make debugging near-impossible when you don't know how users or the AI will act

Air Canada had this thing where um one of their agents predicted or hallucinated a policy um for a refund which was not part of…

47:30
#hallucination#agents#customer-support#risk
Story59:30

The Underwriting Tool That Broke When Users Got Excited

The guests built a system to help bank underwriters pick policies for 30-40 page loan applications, and for three or four months everyone was impressed. Then users got so comfortable they started asking far deeper questions, like throwing the whole document at it and asking what previous underwriters did for similar cases, which required a fundamentally different build. It's a lesson that user behavior evolves and forces recalibration.

  • Underwriting loan applications of 30-40 pages was the painful task being assisted
  • Underwriters reported real time savings for the first three months
  • As users got comfortable they asked harder, unanticipated questions
  • A natural-seeming user request can require a fundamentally different product build
  • User behavior evolves over time and is a trigger to recalibrate the system

post 3 months we realized that they were so excited with the product that they started asking very deep questions that we never anticipated

1:00:00

They would just throw the entire application document at the system and go like for a case that looks like this what did previous underwriters…

1:00:30
#underwriting#user-behavior#calibration#story

Q&A· 1

Q&A41:30

How the Codex Team Really Does Evals (vs "It's All Vibes")

Asked about Claude Code's Boris saying they run on vibes, the OpenAI Codex guest describes a balanced approach. Coding agents are built for customizability, so it's impossible to write evals covering every integration; the team pairs core-protecting evals with heavy customer-signal watching and social-media monitoring. For each new model launch, engineers throw their own hard problems at it.

  • Coding agents are uniquely customizable, so no eval set covers all interactions
  • Evals protect the core; customer behavior and social media catch the rest
  • A/B tests measure whether a new model change actually finds the right mistakes
  • Annoyed users switch the product off, which is itself a signal to watch
  • Each engineer runs custom evals of their own hard problems on every new model

Nah, we don't do evalance on cloud code. It's all vibes.

41:30

if anybody's coming and saying that like my I have this concrete set of evas that I can like bet my life on and then…

44:00
#evals#codex#coding-agents#openai

Tool· 1

Tool1:19:00

Whisper Flow: Transcription That Understands Instructions

In the lightning round the guest recommends Whisper Flow as a favorite recent product. What sets it apart is "conceptual" transcription: used inside a coding tool like Codex it recognizes variable names, and it interprets spoken instructions instead of transcribing them literally — say "add three exclamation marks" and it applies them.

  • Whisper Flow does conceptual transcription, not literal dictation
  • Inside Codex it identifies variables and code context
  • Spoken meta-instructions like 'add three exclamation marks' are executed, not typed out

it's a conceptual transcription tool which means if you go to you know codeex and start using whisfl it starts identifying variables and all of…

1:19:30

you could say something like I'm so excited today add three exclamation marks and it seamlessly switches it adds those three exclamation marks

1:19:30
#tools#transcription#productivity

Takeaway· 2

Takeaway21:30

Reliability, Not Capability, Is Why Enterprises Won't Ship AI

Citing a UC Berkeley / Databricks paper, the guest says roughly three-quarters of enterprises named reliability as their biggest problem, which is why they hesitate to expose end users to AI risk. This also explains why most AI products today are low-autonomy productivity tools rather than end-to-end agents that replace whole workflows.

  • ~74-75% of surveyed enterprises named reliability as their biggest problem
  • Reliability fears keep companies from building customer-facing AI products
  • Productivity use cases dominate because they are lower-autonomy and lower-risk

about 74 or 75% of the enterprises that they had spoken to um their biggest problem was reliability

21:30
#reliability#enterprise#ai-adoption
Takeaway1:04:30

Building Is Cheap Now — Design and Taste Are What's Valuable

The guest argues that with AI, building is cheap and getting cheaper, so the scarce, valuable skill is design: thinking hard about what to build and whether it solves a real pain point. Careers, too, shift after the first few years from execution mechanics toward judgment, taste, and what is uniquely you.

  • Building/implementation is cheap and heading toward ridiculously cheap
  • Design — deciding what to build and whether it solves a pain point — is the expensive part
  • Don't over-obsess with building fast; obsess over the problem and design
  • As execution becomes cheap, careers pivot to taste and judgment

Building is really cheap today. Um design is more expensive.

1:04:30

implementation is going to be ridiculously cheap in the next few years. So really nail down your design, your judgment, your taste and all of…

1:09:00
#taste#design#careers#product-thinking