LLenny's Podcast
← All episodes
11 January 2026

Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google, and Amazon

4Frameworks
15Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster38:00

"Evals" Now Means Five Different Things (Semantic Diffusion)

The guest argues the word "evals" has been butchered into meaning error analysis, expert notes, LLM judges, feedback loops, and even public benchmarks. Invoking Martin Fowler's term "semantic diffusion," they warn that checking LM Arena is not doing evals. The real, agreed-upon goal is building an actionable feedback loop.

  • Data labelers, PMs, and researchers all mean different things by "evals"
  • Martin Fowler's "semantic diffusion": a term gets diluted until its meaning is lost
  • Reading LM Arena or benchmarks is not the same as building your own eval dataset
  • Whether you need an LLM judge depends entirely on context; don't over-prescribe

someone comes up with a term everybody starts butchering it with their own definitions and then you kind of lose the actual definition of it.

39:30

You're not doing eval. That's not eval. Those are model.

41:00
#evals#terminology#llm-judge#benchmarks

Hot Take· 4

Hot Take31:00

"One-Click Agents Are Pure Marketing"

The guest is openly skeptical of vendors promising a one-click agent that delivers gains in days. Enterprise data and infrastructure are too messy for that, and the real work is building a flywheel that improves over time. Meaningful ROI, even with the best data layer, takes four to six months.

  • Vendors promising instant results are ignoring messy enterprise data and tech debt
  • Prefer a company that builds a learning pipeline over one selling out-of-the-box replacement
  • Significant ROI typically takes four to six months even with good infrastructure
  • The edge is the flywheel you build, not being first to deploy an agent

if someone's selling you one click agents, it's it's pure marketing. You don't want to buy into that.

31:30
#agents#enterprise#vendor-hype#flywheel
Hot Take1:01:30

Multi-Agent Systems Are Misunderstood, Not Overhyped

The guest says multi-agent systems are misunderstood rather than overrated. A supervisor agent delegating to sub-agents is a proven, successful pattern. What doesn't work is splitting responsibilities by function and expecting the agents to coordinate on their own in a peer-to-peer "gossip protocol" — current models aren't there yet.

  • Supervisor-with-sub-agents is a genuinely successful pattern
  • Peer-to-peer, self-coordinating agents dividing work by function is the misunderstood part
  • Current model capabilities aren't ready for gossip-protocol multi-agent apps
  • In customer support especially, letting agents talk peer-to-peer makes guardrails unmanageable

if you're building a supervisor agent and there are like sub agents that actually do the work for the super agent supervisor agent that is…

1:02:00

coming with this notion of I'm going to divide the responsibilities based on functionality and somehow uh expect all of that to work together in…

1:02:00
#multi-agent#architecture#agents
Hot Take1:02:30

Coding Agents Are Still Underrated

Despite all the Twitter and Reddit chatter, the guest insists coding agents are underrated. Talk to engineers at ordinary companies, especially outside the Bay Area, and the real-world impact is huge while penetration remains very low. They expect 2025-2026 to create enormous value optimizing these processes.

  • Online hype overstates how far coding-agent adoption has actually spread
  • Outside the Bay Area, real-world penetration is very low despite high potential impact
  • 2025-2026 is expected to be a big value-creation window for coding agents

I still feel coding agents are underrated

1:02:30

talking to an engineer in like any random company uh especially outside of Bay Area you you can see like the amount of impact this…

1:03:00
#coding-agents#adoption#predictions
Hot Take1:12:00

"Pain Is the New Moat"

The guest's favorite concept: in a field with no textbook, the moat isn't being first or having a flashy feature — it's the hard-won knowledge from iterating through pain. Successful companies figured out their non-negotiables and traded them off against model capabilities, and that lived experience is what competitors can't copy.

  • There's no known path or textbook, so knowledge comes from painful iteration
  • Winners identified their non-negotiables and traded them against model capabilities
  • The accumulated org-wide lived experience is the real defensibility
  • Persistence matters because information is now at your fingertips to learn anything

I I like to call it like pain is the new mode

1:12:30

They went through the pain of understanding what are the set of non-negotiable things

1:13:00
#moat#strategy#persistence

Explainer· 3

Explainer07:30

The Two Ways Building AI Products Is Fundamentally Different

The guests argue AI products break from traditional software in two ways most people ignore. First, non-determinism: both the input (users express intent in unlimited natural-language ways) and the output (a probabilistic, black-box LLM) are unpredictable. Second, the agency-control trade-off: giving an agent more autonomy means relinquishing control, so the agent must earn trust before you let it decide.

  • Traditional software has a well-mapped decision engine; AI replaces that layer with a fluid natural-language interface
  • You don't know how the user will behave (input) or how the LLM will respond (output)
  • Every increase in agency is a decrease in your control
  • An agent should build up trust over time before being given more decision-making power

one of them that most people tend to ignore is the non-determinism

08:00

every time you hand over decision-m capabilities or autonomy to agentic systems, you're kind of relinquishing some amount of control on your end

09:30
#ai-products#non-determinism#agents#product-design
Explainer25:30

The Success Triangle: Every Tech Problem Is a People Problem First

The guest frames successful AI companies along three dimensions rather than pure technology: great leaders, good culture, and technical progress. Leaders must relearn intuitions built over a decade, and culture must move from FOMO and replacement fear toward empowerment so subject-matter experts actually cooperate.

  • Three dimensions of success: great leaders, good culture, technical progress
  • Leaders' 10-15 year intuitions must be relearned and they must be vulnerable enough to do it
  • AI adoption almost always needs a top-down push; bottom-up rarely gets leader buy-in
  • A culture of empowerment (not replacement fear) keeps subject-matter experts engaged

Every technology problem is a people problem first.

25:30

great leaders, good culture and technical progress

26:00
#leadership#culture#ai-transformation
Explainer33:30

Evals vs Production Monitoring Is a False Dichotomy

The guest rejects the idea that either evals or production monitoring alone will solve everything. Evals encode your product knowledge into datasets of what should and shouldn't happen; production monitoring surfaces real user behavior through explicit and implicit signals. You need both because each catches errors the other can't.

  • Evals are your product thinking turned into datasets of expected right/wrong behavior
  • Production monitoring reads real usage via explicit and implicit signals
  • Regenerating an answer (not just thumbs-down) is an implicit signal the output missed
  • Evals only catch errors you already anticipated; monitoring finds emerging failure patterns

there's this false dichotomy of like there's either eval is going to solve everything or online monitoring or production monitoring is going to solve everything

33:30
#evals#monitoring#reliability#feedback-loops

Story· 3

Story26:30

The CEO Who Blocks 4-6am Every Day to Catch Up on AI

The guest describes working with the CEO of what is now Rackspace, who blocks 4-6am daily with no meetings just to consume the latest AI content, plus weekend vibe-coding sessions. The point: leaders must get hands-on to rebuild their intuitions and accept they may be the dumbest person in the room.

  • The CEO reserved 4-6am daily, meeting-free, to keep up with AI
  • He curated two or three trusted sources and bounced questions off AI experts
  • Hands-on time is about rebuilding intuition, not implementing things personally

have this block every day in the morning which would say catching up with AI 4 to 6:00 a.m.

26:30

you probably are the dumbest person in the room and you want to learn from everyone

27:00
#leadership#habits#learning
Story47:00

When Air Canada's Agent Invented a Refund Policy

As a cautionary example of over-autonomous agents, the guest cites Air Canada, whose agent hallucinated a refund policy that wasn't in the original playbook, and the company was legally forced to honor it. Alongside their own customer-support product they had to shut down for endless hot-fixing, it motivates building so you don't lose customer trust or let an AI make dangerous decisions.

  • Air Canada's agent hallucinated a refund policy outside its playbook and legal forced compliance
  • The guests once shut down a customer-support product because emerging problems needed endless hot-fixes
  • Big end-to-end agents make debugging near-impossible when you don't know how users or the AI will act

Air Canada had this thing where um one of their agents predicted or hallucinated a policy um for a refund which was not part of…

47:30
#hallucination#agents#customer-support#risk
Story59:30

The Underwriting Tool That Broke When Users Got Excited

The guests built a system to help bank underwriters pick policies for 30-40 page loan applications, and for three or four months everyone was impressed. Then users got so comfortable they started asking far deeper questions, like throwing the whole document at it and asking what previous underwriters did for similar cases, which required a fundamentally different build. It's a lesson that user behavior evolves and forces recalibration.

  • Underwriting loan applications of 30-40 pages was the painful task being assisted
  • Underwriters reported real time savings for the first three months
  • As users got comfortable they asked harder, unanticipated questions
  • A natural-seeming user request can require a fundamentally different product build
  • User behavior evolves over time and is a trigger to recalibrate the system

post 3 months we realized that they were so excited with the product that they started asking very deep questions that we never anticipated

1:00:00

They would just throw the entire application document at the system and go like for a case that looks like this what did previous underwriters…

1:00:30
#underwriting#user-behavior#calibration#story

Q&A· 1

Q&A41:30

How the Codex Team Really Does Evals (vs "It's All Vibes")

Asked about Claude Code's Boris saying they run on vibes, the OpenAI Codex guest describes a balanced approach. Coding agents are built for customizability, so it's impossible to write evals covering every integration; the team pairs core-protecting evals with heavy customer-signal watching and social-media monitoring. For each new model launch, engineers throw their own hard problems at it.

  • Coding agents are uniquely customizable, so no eval set covers all interactions
  • Evals protect the core; customer behavior and social media catch the rest
  • A/B tests measure whether a new model change actually finds the right mistakes
  • Annoyed users switch the product off, which is itself a signal to watch
  • Each engineer runs custom evals of their own hard problems on every new model

Nah, we don't do evalance on cloud code. It's all vibes.

41:30

if anybody's coming and saying that like my I have this concrete set of evas that I can like bet my life on and then…

44:00
#evals#codex#coding-agents#openai

Tool· 1

Tool1:19:00

Whisper Flow: Transcription That Understands Instructions

In the lightning round the guest recommends Whisper Flow as a favorite recent product. What sets it apart is "conceptual" transcription: used inside a coding tool like Codex it recognizes variable names, and it interprets spoken instructions instead of transcribing them literally — say "add three exclamation marks" and it applies them.

  • Whisper Flow does conceptual transcription, not literal dictation
  • Inside Codex it identifies variables and code context
  • Spoken meta-instructions like 'add three exclamation marks' are executed, not typed out

it's a conceptual transcription tool which means if you go to you know codeex and start using whisfl it starts identifying variables and all of…

1:19:30

you could say something like I'm so excited today add three exclamation marks and it seamlessly switches it adds those three exclamation marks

1:19:30
#tools#transcription#productivity

Takeaway· 2

Takeaway21:30

Reliability, Not Capability, Is Why Enterprises Won't Ship AI

Citing a UC Berkeley / Databricks paper, the guest says roughly three-quarters of enterprises named reliability as their biggest problem, which is why they hesitate to expose end users to AI risk. This also explains why most AI products today are low-autonomy productivity tools rather than end-to-end agents that replace whole workflows.

  • ~74-75% of surveyed enterprises named reliability as their biggest problem
  • Reliability fears keep companies from building customer-facing AI products
  • Productivity use cases dominate because they are lower-autonomy and lower-risk

about 74 or 75% of the enterprises that they had spoken to um their biggest problem was reliability

21:30
#reliability#enterprise#ai-adoption
Takeaway1:04:30

Building Is Cheap Now — Design and Taste Are What's Valuable

The guest argues that with AI, building is cheap and getting cheaper, so the scarce, valuable skill is design: thinking hard about what to build and whether it solves a real pain point. Careers, too, shift after the first few years from execution mechanics toward judgment, taste, and what is uniquely you.

  • Building/implementation is cheap and heading toward ridiculously cheap
  • Design — deciding what to build and whether it solves a pain point — is the expensive part
  • Don't over-obsess with building fast; obsess over the problem and design
  • As execution becomes cheap, careers pivot to taste and judgment

Building is really cheap today. Um design is more expensive.

1:04:30

implementation is going to be ridiculously cheap in the next few years. So really nail down your design, your judgment, your taste and all of…

1:09:00
#taste#design#careers#product-thinking