“So most companies that were using 40 should switch to five and five has very different properties.”
GPT-5
By OpenAI
1 recommend/use · 8 sourced episodes
Every sourced reference
Short attributed excerpts only. Timestamps are approximate.
“for a model like GPT5, the number of possible attacks is one followed by a million zeros.”
“Codex has been growing like absolutely explosively since the launch of GPT-5 back in August.”
“fine-tuning is someone taking GPT5 and doing the same sort of thing”
“How does anyone plan a road map when there's just like, okay, JPT5's out?”
“I've been using it for a while.”
“it is also, you know, the smartest um most useful and um fastest Frontier model um that we've ever launched.”
“gbt 5 is like surely going to be extremely useful”
“do you think chat GPT or just say GPT 4 or J GPT 5 GPT 6”
Related frameworks
Adaptive Evaluation Over Static Benchmarks
Measure AI robustness with attackers that learn, not with a frozen dataset of yesterday's attacks
Adapt the Model, Don't Build One
Post-training is the new pre-training — steer an off-the-shelf model to your outcome instead of pre-training your own
Barrels-and-Ammunition Team Design
Staff each team from the gap, not from a fixed PM/EM/designer template
CaMeL Permission Pre-Restriction
Grant an agent only the permissions its stated task needs, decided before it runs
Context Is All You Need Prompting
Treat the model as a brilliant stranger with zero context, and supply what a colleague would already know.
Continuous Calibration, Continuous Development (CCCD)
A CI/CD-style loop for non-deterministic AI: scope, evaluate, deploy, then calibrate against surprises.
Draft-First LLM Augmentation
Never ask the model to do your job — write your version first, then have it improve it.
Emotional Journey Design for Content
Content is predicting reader reactions: hook them, pace the emotion, make people likable.
Evals as Articulating Success
An eval is just a clear spec of ideal behavior — the shared language of AI product work
Evals-Plus-Production-Monitoring Dual Feedback Loop
Reject the false dichotomy: evals catch what you know, production monitoring catches what you don't.
Fake the AI Before You Build It
Never train a model for an MVP — prototype the AI's output and test demand first.
Frustration-Log Micro-Tool Ideation
Beat the idea crisis: log a week of frustrations, then build tiny AI tools to kill them.
Give It Your Hardest Task
Evaluate a serious AI tool on your gnarliest real problem, not a dumbed-down toy.
High Agency, High Urgency Hiring Filter
Hire for two traits only — people who see a problem and go, and people who go now.
Kind and Candid
Reframe candor as an act of kindness so you actually deliver the hard message.
Live in the Future, But Not Too Far
Hold the far-future vision, but land with users where they already work today.
Measure in Hundreds
If your unit of measurement is one hundred attempts, five failures means you have effectively tried zero times.
Mixed-Initiative Contextual Assistance
Surface AI help at the moment it's relevant instead of interrupting with notifications.
Planning in Seasons
Replace rigid roadmaps with secular 'seasons', loose quarterly OKRs, and deliberate slack
Product as Organism: the metabolic loop
Treat an AI product as a living system that ingests signals, tunes on rewards, and improves with every interaction
RAG Data Preparation Over Database Tuning
The biggest RAG quality wins come from preparing data for retrieval, not picking a database.
Randomized Tiered Trial for AI Productivity
Measure whether AI tools help by running a randomized trial split across performance tiers.
Run Toward the Hard Use Cases
Don't disable high-stakes uses to avoid downside — engineer them to be great
Ship-to-Learn: The Emergent-Product Loop
When product properties are emergent, launching is how you discover them
Step-Wise Eval Design for Multi-Step AI Apps
Don't evaluate agents end-to-end; put an eval on every step until you hit coverage.
The Adjacent-Precedent De-Risk Pitch
Win leadership buy-in for a big AI bet by anchoring it to a past bet that already worked.
The Agency-Control Autonomy Ladder
Ship AI in graduated versions, trading human control for machine agency only as trust is earned.
The AI Deployment Risk Triage
Classify any AI deployment into one of three risk tiers before spending a dollar on defense
The AI PM Upskilling Path
Learn the fundamentals, shadow a research scientist an hour a week, and build one model end to end.
The AI Security Vendor Due-Diligence Test
Five questions that expose whether an AI guardrail vendor is selling real protection or theater
The AI Success Triangle
Successful AI adoption is a people problem first: great leaders, good culture, and technical progress.
The Angry-God Containment Lens
Assume the AI is a malicious agent trying to hurt you, then engineer so it structurally can't
The Better Tool, Same Problems Lens
Every model leap gets normalised within months — build for the boring future, not the euphoric one.
The Constrained-Resource Headcount Test
When the bottleneck isn't people, each new hire is a net productivity loss unless they uplevel everyone.
The Enterprise AI Adoption Ladder
Three sequential stages that separate companies winning with AI from those spinning their wheels
The Eval ROI Decision
Build evals where failure is catastrophic or you must win; vibe-check the rest.
The Four Challenges of AI Product Management
Uncertainty, pivots, data scarcity and a broken promo path — the four taxes of the AI PM role.
The Maximally Accelerated Question
A forcing question that separates critical path from what can wait
The Model Launch Bar
With probabilistic products, the PM — not the scientist — decides what accuracy is good enough to ship.
The Platform Encroachment Test
Build where the platform's mission says it will never go — the general layer is not yours to own.
The Proxy Goal Ladder
Every metric you chase is a proxy — climb the ladder to the mission before you optimise it.
The Shiny Object Trap (Problem-First AI)
A regular PM ships the right product; an AI PM solves the right problem.
The Teammate Onboarding Model for AI Agents
Adopt a coding agent the way you'd onboard a new intern — pair first, then delegate.
Treat Your Course Like a Product
Hypothesise the audience, interview them, iterate the ICP, and run three weeks — not one.
Two-Mode Prioritization for AI Products
Prioritize backward from model magic AND forward from customer needs
Two-Question New Technology Adoption Test
Before adopting any new AI tool, ask: how big is the gain, and how painful is the exit?
Unblock the Review Bottleneck
The limiting factor on AI productivity is human review speed — engineer the agent to validate its own work.
What Actually Improves AI Apps
Stop chasing AI news and vector DBs; the real levers are users, data, and prompts.
Win on the Platform, Not the Features
Durable products win on invisible infrastructure — reliability, reach, privacy — not on the feature list
People in these episodes
Related resources
- A Fire Upon the Deep
- Airbnb
- Airtable
- Alexa
- Amazon
- Amplitude
- Anthropic
- App Store
- Apple
- Apple Podcasts
- Arcade
- Artificial Analysis
- arXiv
- Atlas
- AutoML
- Booking.com
- Boz's podcast
- Brilliant Smart Home System
- caffeinate
- Canva
- CareerFoundry
- ChatGPT
- Chime
- Clair Obscur: Expedition 33
- Claude Code
- Codex
- Coding Dojo
- Comet
- Conductor
- Coursera
- Cursor
- DALL-E
- Databricks
- Datadog
- DeepMind
- DeWalt PowerPack
- Discord
- Disney
- Don't Write That Jailbreak Paper
- Dropbox
- dscout
- DX
- Eppo
- F1
- Figma
- Fin
- For All Mankind
- From Third World to First
- Gemini
- General Assembly
- GitHub
- GitHub Copilot
- Google Cloud
- Google Docs
- Google Forms
- Google Glass
- Google I/O
- Google Maps
- Google Sheets
- Goose
- GPT Store
- GPT-3
- GPT-4
- GPT-4o
- GPTs
- Gran Turismo
- Hack Prompt and Learn Prompting AI Security Course
- Harvey
- Her
- Hex
- High Output Management
- Hugging Face
- Inspired
- Instacart
- Interpret
- Introduction to Artificial Intelligence
- iPod
- Jujutsu Kaisen
- Justworks
- Lenny's Newsletter
- Lenny's Podcast
- Lenny's Podcast website
- Lennybot
- lennyspodcast.com
- Lensa
- Linear
- LM Arena
- Manta Sleep Mask
- Manus
- Marginal Revolution
- Maven
- Messenger
- Meta
- Metronome
- Microsoft
- Microsoft PowerPoint
- Microsoft Windows
- Microsoft Word
- Midjourney
- Model Context Protocol
- NASA
- Netflix
- Notion
- Nvidia
- NVIDIA NeMo
- o3
- OpenAI
- OpenAI API
- OpenAI Prompt Engineering Guide
- Orkes
- Pando
- Persona
- PostHog
- Python
- QuickBooks
- Quip
- r/Codex
- Radical Candor
- Raycast
- React
- Remembrance of Earth's Past
- Repello AI
- Runway
- Salesforce
- SAP
- Scale AI
- Sentry
- ServiceNow
- Severance
- Sierra
- Silicon Valley
- Slack
- Sora
- Sora Android app
- Spark
- Spotify
- SQL
- Stanford University
- Stranger Things
- Tesla
- The Culture
- The Design of Everyday Things
- The Download
- The Lord of the Rings
- The One World Schoolhouse
- The Selfish Gene
- The White Lotus
- TikTok
- TLDR
- Universal Primer
- Van Westendorp Price Sensitivity Meter
- Vanta
- Visual Electric
- Visual Studio Code
- WAOAW Sleep Mask
- Waymo
- Westworld
- When Breath Becomes Air
- Whimsical
- Why We Sleep
- Wikipedia
- Wispr Flow
- X
- You Look Like a Thing and I Love You
- YouTube
- Zapier
Spot an error or want this page removed? Request a correction or removal.