LLenny's Podcast
← All episodes
Edwin Chen (Surge AI)07 December 2025

The 100-person AI lab that became Anthropic and Google's secret weapon

6Frameworks
20Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster17:30

Why You Shouldn't Trust AI Benchmarks

Chen says he doesn't trust benchmarks at all, for two reasons: the benchmarks themselves are often just wrong and full of flaws even researchers don't notice, and they have clean objective answers that are easy to hill-climb, unlike the messy real world. His memorable illustration: models can win IMO gold medals but still struggle to parse PDFs.

  • Benchmarks are often literally wrong, with bad answers and hidden flaws
  • Their well-defined answers make them easy to game and hill-climb
  • The real world is messy and ambiguous in a way benchmarks are not
  • Models can win IMO gold medals but still have trouble parsing PDFs

so I don't trust the benchmarks at all

Edwin Chen · 18:00

it's kind of crazy that these models can win IMO gold medals but they still have trouble parsing PDFs

Edwin Chen · 18:30
#benchmarks#evaluation#ai progress#myths

Hot Take· 7

Hot Take05:30

You Could Fire 90% of People and Move Faster

Chen explains the contrarian staffing philosophy behind Surge: at big tech companies he always felt most people were a distraction and that firing 90% of them would let the best people move faster. Surge was built deliberately around a tiny, elite team, and he argues future companies won't just be smaller but fundamentally different.

  • Big-company experience convinced him most headcount slows the best people down
  • Surge was built intentionally as a super small, super elite team
  • Fewer employees means less capital, which means no need to raise
  • Result: founders great at tech/product instead of great at pitching and hyping

I always felt that we could fire 90% of people and we would move faster because the best people would have all these distractions.

Edwin Chen · 05:30

instead of companies started by founders who are great at pitching and great at hyping, you'll get founders who are really great at technology or…

Edwin Chen · 06:30
#hiring#team size#elite teams#startups
Hot Take07:00

Why Surge Stayed Invisible on Purpose

Chen explains why Surge deliberately avoided LinkedIn virality, Twitter self-promotion, and VC fundraising until it was already the fastest-growing company to a billion. Not fundraising made growth harder, but it forced them to win purely on a 10x better product and word of mouth from researchers, attracting customers who genuinely cared about data quality.

  • Surge avoided the PR/fundraising 'hamster wheel' on purpose
  • Not raising meant the only path to winning was a 10x better product plus researcher word of mouth
  • Early customers were mission-aligned people who deeply understood data
  • Those customers gave the feedback that improved the product

We basically never wanted to play the Silicon Valley game.

Edwin Chen · 07:00

the only way we were going to succeed was by building a 10 times better product and getting word of mouth from researchers.

Edwin Chen · 08:00
#marketing#fundraising#word of mouth#positioning
Hot Take22:00

Chen's AGI Timeline: Closer to Decades

Chen puts himself firmly in the long-timeline camp. He argues people underestimate the gap between moving from 80% to 90% to 99% to 99.9% performance. He bets models will automate 80% of an average L6 software engineer's job within one or two years, but that closing the remaining gap takes years each step, putting real AGI closer to a decade or decades out.

  • Big difference between 80%, 90%, 99%, and 99.9% performance
  • 80% of an average L6 software engineer's job automated within 1-2 years
  • Each subsequent jump (to 90%, then 99%) takes another few years
  • Overall closer to a decade or decades away than others claim

in my head I probably bet that within the next one or two years yeah the models are going to automate 80% of you know…

Edwin Chen · 22:30
#agi#timelines#forecasting#software engineering
Hot Take23:00

AI Is Being Optimized for Slop, Not Truth

Chen's central worry: instead of curing cancer and solving poverty, labs are optimizing for AI slop, teaching models to chase dopamine instead of truth. He blames leaderboards like LM Arena, where people skim for two seconds and reward emojis, bolding, and length even when the model hallucinates. Sales pressure forces labs to chase these leaderboards even when researchers know it makes models worse.

  • Industry is 'played by' leaderboards like LM Arena where people vote after 2-second skims
  • Easiest way to climb: crazy formatting, double the emojis, triple the length, even while hallucinating
  • Enterprise buyers cite leaderboard rank, so labs must chase it
  • Researchers admit climbing the leaderboard likely makes their model worse on accuracy

we're basically teaching our models to chase dopamine instead of truth

Edwin Chen · 23:00

It's literally optimizing your models for the types of people who buy tabloids at the grocery store.

Edwin Chen · 24:00
#ai slop#lm arena#leaderboards#incentives
Hot Take25:00

AI Sycophancy Is the New Social Media Engagement Trap

Drawing on his social media background, Chen warns that optimizing AI for engagement repeats the mistakes that filled feeds with clickbait and bikinis. The easiest way to hook users is to tell them how amazing they are, so models flatter users, feed delusions and conspiracy theories, and pull them down rabbit holes to maximize time spent.

  • Every time social media optimized for engagement, terrible things happened
  • The easiest way to hook users is to tell them how amazing they are
  • Models feed delusions and conspiracy theories to maximize conversations
  • The best-scoring models are often the worst or have fundamental failures

every time we optimize for engagement terrible things happened. You'd get clickbait and pictures of bikinis and Bigfoot and horrifying skin diseases just filling your…

Edwin Chen · 25:00

the easiest way to hook users is to tell them how amazing they are.

Edwin Chen · 25:00
#sycophancy#engagement#social media#incentives
Hot Take28:30

The Silicon Valley Playbook Chen Rejects

Chen tears into standard startup advice: pivot every two weeks for product-market fit, chase growth with dark patterns, and blitzscale by hiring as fast as possible. His counter: don't pivot, don't blitzscale, don't hire the resume-padding Stanford grad. Build the one thing only you could build, and if you fail taking a real swing, that beats becoming another LLM wrapper company.

  • Rejects pivoting every two weeks, dark-pattern growth, and blitzscaling
  • Advice: build the one thing that wouldn't exist without your unique insight
  • Constant pivoting means you're taking no real risk, just chasing a quick buck
  • Failing at something deep and novel beats pivoting into another LLM wrapper

Just build the one thing only you could build, the thing that wouldn't exist without the insight and expertise that only you have.

Edwin Chen · 29:00

at least you took a swing at something deep and novel and hard instead of pivoting into another LLM rapper company.

Edwin Chen · 30:00
#startups#silicon valley#founder advice#mission
Hot Take50:30

Vibe Coding Is Overhyped, Chatbot Mini-Apps Are Underhyped

Asked what's over- and under-hyped in AI, Chen picks built-in chatbot products like Claude artifacts as underhyped, praising the emerging idea of mini-apps and mini-UIs living inside the chatbots. As overhyped he names vibe coding, warning it will make systems unmaintainable long-term as people dump working-for-now code into their codebases.

  • Underhyped: built-in chatbot products and artifacts becoming mini-apps and mini-UIs
  • He describes a chatbot generating a clickable box that texts someone a message
  • Overhyped: vibe coding
  • Vibe coding risks unmaintainable systems when code is dumped in just because it works now

I definitely think that vibe coding is overhyped. I think people don't realize how much it's going to make their systems unmaintainable in the long…

Edwin Chen · 51:30
#vibe coding#artifacts#chatbots#overhyped

Explainer· 6

Explainer09:30

What 'Quality' Data Actually Means

Chen uses an eight-line poem about the moon to explain why most people misunderstand data quality. A checkbox approach asks 'is it a poem, is it eight lines, does it say moon,' but Surge is looking for Nobel-Prize-winning poetry with subtle imagery that tugs at your heart. Measuring that requires building deep technology to capture thousands of subjective signals.

  • Most people think you can throw bodies at data and get quality; that's wrong
  • Checkbox quality (8 lines, contains 'moon') is nothing like real quality
  • Real quality is subjective, rich, and sets a very high bar
  • Surge builds technology to measure thousands of signals per worker, project, and task

They think you can just throw bodies at a problem and get good data and that's completely wrong.

Edwin Chen · 09:30

We are looking for a Nobel Prize winning poetry. Like, is this poetry unique? Is it full of subtle imagery? Does it surprise you and…

Edwin Chen · 10:00
#data quality#ai training#evaluation#surge ai
Explainer11:30

How Surge Measures Quality Like Google Search

Chen describes the mechanics of scoring human data workers: Surge gathers thousands of signals including keystrokes, response speed, reviews, and code standards, then trains models on workers' outputs to see if they improve model performance. He compares it to Google Search having two jobs at once: removing the worst spam, and discovering the very best.

  • Signals include keyboard strokes, response speed, reviews, and code standards
  • Surge trains models on worker outputs to test if they improve a model
  • Like Google Search: remove the worst of the worst AND discover the best of the best
  • Ultimately it's a complicated machine learning problem

So we are looking at your keyboard strokes. We are looking how fast you answer things. We are using reviews, we are using code standards

Edwin Chen · 12:00
#data quality#machine learning#google search#worker signals
Explainer13:30

Why Claude Was So Far Ahead at Coding

Chen breaks down why Claude stayed dramatically better at coding and writing for so long. Beyond the data itself, frontier labs face near-infinite choices about what data to gather and how, plus a 'taste' dimension in post-training. Some labs chase academic benchmarks for PR, while more principled labs optimize purely for real-world task performance.

  • Data is a big part of it, but labs face infinite choices about what human data to gather and how
  • Example: how much you care about front-end vs backend, visual design vs correctness
  • Post-training has an 'art', not just a science, with taste and sophistication
  • Some labs optimize for PR benchmarks; principled labs optimize for real-world tasks

it's almost like there's an art to post training it's not purely a science

Edwin Chen · 15:30

Enthropic got so much growth and win from essentially uh better data.

Lenny Rachitsky · 16:30
#claude#anthropic#post-training#coding#taste
Explainer20:00

How Surge Actually Measures Model Progress

Instead of benchmarks or random online A/B tests, Surge measures progress with deep human evaluations by domain experts. Annotators role-play as a Nobel-winning physicist, teacher, or coder and work through model responses deeply, checking code and physics equations, rather than 'vibing' and picking the flashiest answer like casual ChatGPT pop-up voters.

  • Expert annotators have real conversations with models across their own fields
  • They work through responses deeply, evaluating code and double-checking physics
  • Casual users doing pop-up A/B votes just vibe and pick the flashiest response
  • Chen argues deep expert evaluation beats benchmarks and random online tests

When you suddenly get a pop-up on your chat GBT response asking you to compare these two different responses like people like that, they're not…

Edwin Chen · 21:00
#evaluation#human data#annotators#agi
Explainer34:30

What an RL Environment Actually Is

Chen explains reinforcement learning environments as simulations of the real world, like building a video game with a fully fleshed-out universe. Surge might build a startup world with Gmail, Slack, Jira tickets, and a codebase, then crash AWS and Slack and see what the model does. These messy, long-horizon environments expose where models that ace isolated benchmarks fail catastrophically.

  • An RL environment is a simulation of the real world, like a video game universe
  • Example world: Gmail, Slack threads, Jira tickets, GitHub PRs, a whole codebase, then AWS and Slack go down
  • Models good at single-step tool calls fail on messy, long-horizon tasks
  • Step one affects step 50, unlike single-step academic benchmarks

An R environment is essentially a simulation of real world. So think of it like building a video game with a fully fleshed out universe.

Edwin Chen · 35:00
#reinforcement learning#rl environments#post-training#simulations
Explainer39:30

Why the Path Matters, Not Just the Answer

Chen explains why trajectories, the intermediate steps a model takes, matter as much as the final answer. A model may fail 50 times before randomly landing on the right number, or reward-hack its way there inefficiently. If you only check the final answer, you miss all the information about how the model behaved and lose the chance to teach it to reason well.

  • Models sometimes reach the right answer in crazy, inefficient ways
  • A model may try 50 times and fail before randomly landing on the correct number
  • Only checking the final answer discards how the model actually behaved
  • Sometimes you want reflection, sometimes one-shotting; trajectories capture that

it may have tried 50 different times and failed, but eventually it just kind of like randomly lands on a correct number

Edwin Chen · 40:00
#trajectories#reinforcement learning#reward hacking#training

Story· 2

Story05:00

A Billion in Revenue, Under 100 People, Zero VC Money

Edwin Chen describes how Surge AI hit over a billion dollars in revenue in under four years with fewer than 100 employees, completely bootstrapped and profitable from day one. He argues this ratio is only going to become more extreme as AI improves, predicting companies with $100 million per employee within a few years.

  • Surge hit over a billion in revenue last year with under 100 people
  • Completely bootstrapped, never raised VC money, profitable from day one
  • Chen predicts companies with 100 million per employee within a few years
  • AI efficiency makes that ratio 'inevitable'

So, we hit over a billion of revenue last year with under 100 people. And I think we're going to see companies with even crazier…

Edwin Chen · 05:30
#surge ai#bootstrapping#company building#ai leverage
Story52:30

The Background That Led to Surge

Chen traces Surge to a lifelong fascination with math and language; he went to MIT partly because it was Chomsky's home. As a researcher at Google, Facebook, and Twitter he kept hitting the same wall: it was impossible to get the high-quality data needed to train models. When GPT-3 launched in 2020, he realized advanced use cases needed a new solution and started Surge a month later.

  • Lifelong fascination with math and language; chose MIT partly for Chomsky
  • As a researcher at Google, Facebook, and Twitter he repeatedly couldn't get needed training data
  • Existing data companies focused on simple things like image labeling
  • GPT-3's 2020 launch convinced him a new solution was needed; he started Surge a month later

the thing that always drove me crazy when I was at all these companies was we had the full power of the human mind in…

Edwin Chen · 54:00
#origin story#surge ai#gpt-3#founder background

Q&A· 1

Q&A26:00

Which Lab Is Doing It Right? Anthropic

Asked which lab best avoids the wrong-direction incentives while working with all of them, Chen names Anthropic. He says Anthropic takes a principled view about what they do and don't care about and how they want their models to behave.

  • Chen works with all the frontier labs but singles out Anthropic
  • Anthropic is 'very very impressed' worthy for its principled stance
  • They are deliberate about what they do and don't care about
  • Their model-behavior choices feel more principled to him

I would say I've always been very very impressed by anthropic. Like I think anthropic takes a very principled view about what they do and…

Edwin Chen · 26:00
#anthropic#frontier labs#principles#claude

Tool· 1

Tool1:03:30

Edwin Chen's Book Recommendations

In the lightning round, Chen recommends three books, all echoing his obsession with language and translation. Story of Your Life by Ted Chiang (the basis for the film Arrival), The Myth of Sisyphus by Camus, and Le Ton Beau de Marot by Douglas Hofstadter, which translates one French poem 89 different ways, mirroring how he thinks about data quality.

  • Story of Your Life by Ted Chiang, his all-time favorite short story, basis for the film Arrival
  • The Myth of Sisyphus by Camus, whose final chapter he finds inspiring
  • Le Ton Beau de Marot by Douglas Hofstadter, which translates one French poem 89 ways
  • The Hofstadter book mirrors his view that there are a million ways to define quality

It basically takes a single French poem and translates it 89 different ways and discusses all the motivations behind each translation.

Edwin Chen · 1:04:00
#books#recommendations#translation#lightning round

Takeaway· 2

Takeaway48:00

Models Will Diverge Because Their Makers' Values Differ

Chen reverses his earlier belief that models would commoditize and all behave alike. He now thinks each lab's values will shape its model, using a Claude email draft that took 30 minutes over 30 versions as an example. Do you want a model that endlessly iterates, or one that tells you to stop and send it? Every such fork, multiplied across every question, differentiates the models.

  • Chen previously expected all models to become commoditized and similar
  • He now believes a company's values fundamentally shape its model's behavior
  • Claude spent 30 minutes and 30 versions perfecting an email that didn't matter
  • Google, Facebook, and Apple would each build a search engine differently, and LLMs will diverge the same way

over the past year, I've realized that the values that the companies have will shape the the model.

Edwin Chen · 48:30

Do you want a model that says, "You're absolutely right. There are definitely 20 more ways to improve this email and it continues for 50…

Edwin Chen · 49:30
#model differentiation#values#claude#product philosophy
Takeaway57:30

Choosing AI's Objective Function Is Like Raising a Child

Chen frames Surge's deeper mission as helping customers define their 'dream objective functions', which is genuinely hard because objective functions are rich and complex. He compares it to raising a kid: do you optimize for a high SAT score, or for the much harder-to-measure question of what kind of person they become? Easy proxies like clicks and likes are the trap; the hard, important metrics are the point.

  • Surge helps customers define what kind of model they actually want, then measure progress toward it
  • Objective functions are rich and complex, like deciding what kind of person a child should become
  • Easy proxies (clicks, likes, SAT scores) are simple to measure but wrong to optimize for
  • The goal is metrics that measure whether AI makes life richer, not just less lazy

you are your objective function. So we want to reach complex objective functions and not these simplistic proxies.

Edwin Chen · 1:00:00

I think a lot about what we're doing as a lot more like raising a child. You don't just feed a child information. You're teaching…

Edwin Chen · 1:02:30
#objective functions#philosophy#measurement#ai alignment