LLenny's Podcast
← All episodes
Jason Droege09 October 2025

First interview with Scale AI’s CEO: $14B Meta deal, what’s working in enterprise AI, and what frontier labs are building next

7Frameworks
15Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 2

Myth Buster13:30

The 'it's all cheap generalist labor' story is bogus

Droege pushes back hard on competitor positioning that Scale still relies on low-cost generalist labelers. The industry moved from basic preference ranking (is this short story better than that one?) to expert tasks that take hours and require PhDs and professionals — like building a full website or explaining nuanced cancer topics to a model. He backs it with the makeup of Scale's expert network.

  • 18 months ago a typical task was ranking which short story was better; now a task is a top web developer building an entire site or a PhD explaining cancer nuance
  • Tasks now take hours and require professionals and PhDs
  • 80% of Scale's expert network has a bachelor's degree or greater; about 15% have a PhD
  • PhDs on the network earn significant money contributing expertise to models
  • Scale often flags model weaknesses to labs and offers experts to fix them

I think I think the the the current um positioning out there from competitors is just bogus.

Jason Droege · 13:30

80% of the people uh that we have in our expert network have a bachelor's degree or greater, which is very contrary to some of…

Jason Droege · 16:00
#data labeling#experts#evals#training data
Myth Buster37:30

Why enterprise AI pilots 'fail' — and the real timeline to value

On the MIT study and headlines that 95% of enterprise AI pilots fail, Droege calls the number a bit of clickbait distorted by a denominator effect — it's so easy to spin up a project that failures pile up. With a quality partner or experienced engineers putting in months (not the minutes shown in demos), the impact is real. But robustly automating an important process takes six to twelve months, longer than what people are selling.

  • The 95% failure number is inflated by a denominator effect — projects are trivially easy to start
  • Reliability behaves like data-center uptime: each additional 'nine' is an order-of-magnitude more investment
  • Real results require legal, policy, regulatory approval, change management, and an accuracy everyone is comfortable with
  • Robustly automating an important process takes six to twelve months
  • When it works the impact is huge, but the time to get there is longer than vendors claim — 'easy to learn, hard to master'

so I don't necessarily know that like the 95% number I think is a bit of like clickbait in a way um it tells the…

Jason Droege · 39:00

these things take six to 12 months to get them truly, you know, robust enough where like an important process can be automated.

Jason Droege · 39:30
#enterprise ai#pilots#adoption#roi

Hot Take· 4

Hot Take28:30

Why humans keep training AI: 'a history of new beginnings'

Asked how long labs will need human experts, Droege frames data labeling as a history of new beginnings — old needs fade (autonomous vehicles) while new ones constantly appear. The point where no new human skill or knowledge matters to models feels far out. On the broader white-collar-apocalypse fear, he sides with practicality: change will come, but humans are highly adaptable.

  • Data labeling is a history of new beginnings — needs shift as models improve
  • The point where no new human knowledge matters to models feels far off
  • Scale continuously finds new needs, sometimes surfacing expertise its existing contributors didn't know was useful
  • Human-in-the-loop is a personal belief, not just a business one — these systems need to work for us
  • He is skeptical the labor transformation happens in the next one to two years; humans adapt

the history of data labeling is a history of new beginnings

Jason Droege · 28:30

I think what we're underestimating in all of the doom and gloom is we believe in human adaptability.

Jason Droege · 30:30
#future of work#labor#human-in-the-loop#adaptability
Hot Take35:30

The next two years: from models knowing things to models doing things

Droege's read on where models head next: knowledge is already robust, so the frontier shifts to action — navigating a Salesforce instance, a healthcare system, even a weather app, and making decisions for you. We're at the very beginning of that, which is why forecasts vary so widely. His guess is that within two to three years the technology gets close enough to push change-management and policy questions to the front.

  • Model knowledge is already quite robust; the open question is what it can do for you
  • Agentic action is where RL environments come into play
  • We're at the beginning, so trajectories and speculation vary widely
  • Adoption may become a human and policy issue, not a technological one — some people still lack an email address
  • His guess: in two to three years the tech gets close enough to force change-management and policy debates

the general trend right now is going from models knowing things to models doing things

Jason Droege · 35:30
#agents#ai forecast#adoption#policy
Hot Take48:30

The founder's real question: why am I lucky enough to have this insight?

Droege's philosophy on starting anything: you're hunting for alpha, so if your research is shaped by what everyone else is saying you won't have an independent insight. The bar is asking why you, among a million smart entrepreneurs, would have an insight others don't — and why you'd want to work on it for five to ten years. He pairs contrarian insight with two success factors: a founder who is a force of nature through endless pivots, and knowing which business models actually work.

  • Entrepreneurship is a search for alpha — independent insight others don't have
  • Ask why you're lucky enough to have this insight, and why you'd work on it for 5-10 years
  • Don't fall in love with your ideas; be willing to throw away who you've been for the mission
  • Most important success factor: a founder who's a force of nature through years of pivots
  • Second: knowing good vs. bad business models — marketplaces, sticky recurring-revenue SaaS, network effects, more valuable at scale

why am I in the position where I likely have an insight that others do not?

Jason Droege · 49:30

the founder is just a force of nature over a long duration of time because you're going to have to pivot.

Jason Droege · 51:00
#founders#entrepreneurship#business models#contrarian
Hot Take65:30

Not losing is a precursor to winning

Against the tech consensus to 'just go for it,' Droege argues the best entrepreneurs weigh the risk profile of decisions and make asymmetrically positive bets. Survival is part of the game — most people give up before their timing, insight, or product clicks, and life in tech can flip from dog to hero fast. In a hype cycle, the temptation to go for it more and more risks compromising the enterprise you need to keep alive to serve customers.

  • The 'just go for it' narrative is consensus, largely driven by investors
  • Best entrepreneurs look at the risk profile and make asymmetrically positive decisions
  • Survival is part of the game; most give up before timing and insight align
  • You can go from dog to hero fast, but only if you survive long enough
  • In a hype cycle, don't put the enterprise in a position that could compromise it — take risk, but calculate it

the best entrepreneurs and the best business owners look at the risk profile of the decisions that they're making and they try to make asymmetrically…

Jason Droege · 65:30

Survival is just part of the game and most people just give up before they could get their timing right

Jason Droege · 66:00
#risk#founders#survival#decision-making

Explainer· 3

Explainer10:30

What the $14B Meta deal actually was, and what Scale looks like now

Droege sets the record straight on the Meta transaction that confused the market. Meta invested over $14B for 49% non-voting stock with no new board seat, no preferential access to data, and only about 15 people leaving in the deal. Scale remains fully independent with roughly two unicorn-scale businesses and has grown every month since the deal closed.

  • Meta invested a little over $14B for 49% non-voting stock; governance and the board are largely unchanged
  • No preferential access or relationship — same privacy and data security as before
  • Only about 15 people moved to Meta in the transaction
  • Scale has two major businesses, each doing hundreds of millions in revenue, and has grown every month since the deal

The transaction was uh Meta invested a little bit over 14 billion to get 49% of the company non- voting stock.

Jason Droege · 10:30

only about 15 people went over in the transaction.

Jason Droege · 11:30
#scale ai#meta#m&a#enterprise ai
Explainer19:00

What RL environments are and why generalizability is the whole game

Droege explains RL environments as sandboxes where AI agents practice accomplishing a goal — like navigating a highly configurable Salesforce instance and knowing when to escalate to a human. Because the permutations of environments, configurations, and goals are effectively endless, the labs' real need is data generalizable enough that they don't have to collect trillions of combinations.

  • RL environments are sandboxes for agents to learn how to accomplish a goal
  • Example: an agent navigating a Salesforce instance, recognizing data and config, and escalating to a human at low confidence
  • The number of environments and goals within each is enormous — permutations are endless
  • The value of the data rises with how generalizable it is across use cases
  • Scale's job is supplying the most generalizable, valuable data so labs avoid collecting every combination

there's these things called RL environments that effectively are sandboxes for AI agents to play in to accomplish a goal so that they can learn…

Jason Droege · 19:00
#reinforcement learning#agents#rl environments#salesforce
Explainer31:30

Evals, simply: 'what does good look like?'

Droege reduces evals to a single question — what does good look like? Especially for enterprise and government customers, most of the expert work is eval: someone has to establish the benchmark for good. In the healthcare case, the doctor effectively creates evals defining what a report should surface. He distinguishes 'good' from 'correct' because these are probabilistic systems making judgment calls.

  • Most enterprise and government expert work is eval
  • An eval establishes the benchmark for what good looks like
  • AI shines where a human process is only 10-20% accurate and gets to 50-80%; it struggles to close the last 2% of a 98%-accurate process
  • He says 'good' rather than 'correct' deliberately — these are probabilistic systems
  • Many tasks are asking the system for the best recommendation given current information, like you'd ask a person

somebody's got to establish the benchmark for like what good looks like. That's the simple way to think about eval. What does good look like?

Jason Droege · 31:30
#evals#benchmarks#enterprise ai#accuracy

Story· 3

Story23:30

How Scale's healthcare tool caught an allergy a doctor would have missed

Droege shares a concrete applications-side example. A health system's specialists see rare cases with huge backlogs, and diagnosing them can require reading 200-300 pages of documentation in mixed formats. Scale built a tool that reads the document and surfaces the top 5-10 things to consider — in one case flagging a non-obvious allergy that conflicted with the medication about to be prescribed. It illustrates why labeling is now moving inside enterprises.

  • Rare-case specialists face big backlogs and want accurate day-one diagnoses to prevent revisits
  • Diagnosis can require reading 200-300 pages of mixed-format documentation
  • Scale's tool reads the document and points out the top 5-10 things to consider
  • It caught a non-obvious allergy that conflicted with a planned medication
  • Off-the-shelf models plus RAG and fine-tuning only go so far — labeling is moving into enterprises and governments

the doctor really needs to read two to 300 pages of documentation

Jason Droege · 24:30

we picked up on an allergy that a patient had that would not have been obvious from reading the document

Jason Droege · 25:00
#healthcare#enterprise ai#labeling#case study
Story44:00

How Uber Eats reverse-engineered restaurant economics with a scale

When restaurants wouldn't share unit economics, Droege's team built their own ground truth — ordering food and weighing the ham, cheese, and bread to compose an independent view of ingredient vs. labor cost. That insight (roughly 20-30% ingredients, 20-30% labor, ~10% real estate) let them confidently pitch a 30% take rate against a true clearing price near 25%. The deeper lesson: understand a customer's real incentives and urgency better than they articulate it.

  • Restaurants were suspicious and wouldn't share unit economics
  • The team ordered food and weighed ingredients to build their own ground truth
  • Rough structure: 20-30% ingredients, 20-30% labor, ~10% real estate
  • They pitched 30% take; the real clearing price landed near 25%
  • The winning frame was incremental demand at high incremental gross margin, not satisfying any one party 100%
  • Biggest miss founders make: ignoring the urgency of the buyer, not just the value provided

we just matched up like how much does the ham weigh, how much does the cheese weigh, how much does the bread weigh, how many…

Jason Droege · 44:30

the biggest thing people miss when they're building new products is the urgency of the buyer part of it.

Jason Droege · 48:00
#uber eats#unit economics#customer research#marketplaces
Story57:30

How saying no to McDonald's won Uber Eats a better deal

McDonald's approached Uber Eats wanting to do delivery, and Droege said no — his vision was helping local independents, not chains, and chains scared everyone off over basket-size economics. He pushed them off for four or five months until his team called him insane. That reluctance helped land an exclusive relationship and an insane number of customers, and the McDonald's launch went global in about six months while Uber Eats was under two years old.

  • Droege initially refused McDonald's because his vision was empowering local independents
  • Chains weren't really on delivery networks yet due to basket-size sensitivity
  • He pushed McDonald's off for four or five months over his team's objections
  • The reluctance helped secure an exclusive deal and a huge customer influx
  • His approach to the tough economics: 'figure it out' — reduce delivery radius, mark up some items
  • They took McDonald's global in about six months while Uber Eats was under two years old

McDonald's actually approached us and they said, "Hey, we'd love to do food delivery with you." And I'm and I said no.

Jason Droege · 58:00

I push them off for like four or five months until my team is like, "You're insane.

Jason Droege · 58:30
#uber eats#mcdonalds#deals#negotiation

Tool· 1

Tool1:12:00

Droege's daily AI habit: a voice-mode tutor on the way to work

Coming from a consumer background into a fast-moving technical space, Droege's most impactful use of AI is as a tutor — turning on voice mode and talking to it on his commute to stay on top of new concepts his researchers don't always have time to explain. His second use: feeding internal documents to AI and asking what's most important, which he says is shockingly good at cutting through the organizational broadcast problem.

  • He uses AI as a tutor, turning on voice mode and talking to it on his way into work
  • It fills the gap when busy researchers can't explain fast-moving new concepts
  • Second use: ask AI for the most important thing in an internal document, then read and double-check
  • He finds it shockingly good at surfacing what matters amid organizational noise
  • He also applies it to legal documents to understand what's for or against him

I use it as a tutor like like I turn on voice mode and talk to it on my way into work.

Jason Droege · 1:13:00
#ai tools#productivity#voice mode#learning

Takeaway· 2

Takeaway60:30

Gross margin as a coarse litmus test for whether an idea is any good

Droege uses gross margin as a quick, imperfect filter on new businesses: if you can't mark something up much, how much value are you really adding? His favorite move when someone pitches a 40% gross margin is to ask why it isn't 60% — which immediately surfaces the real problem, usually a cheaper alternative that will compress your margin faster than you think. The deeper question is always why can't someone else do this in two years.

  • High gross margins plus healthy churn curves are a healthy sign for a business
  • If you can't mark it up, you're probably not adding much value
  • His trick: ask why a proposed 40% margin isn't 60% — it short-circuits to the real problem
  • An existing low-margin alternative signals your margin will compress quickly
  • The follow-up question is always 'why can't someone else do this in two years?' — if they can, expect margin compression
  • It's a coarse instrument, not perfect — some great businesses (Costco, Walmart) run low margins by design

high gross margins uh, combined with healthy churn curves are a very healthy sign for the business.

Jason Droege · 60:30

someone comes up with an idea and they go we can get into this business and I think we can charge this and it'll get…

Jason Droege · 61:00
#gross margin#business models#unit economics#founders
Takeaway1:09:30

Build the right team, and interview for just three things

Droege argues building the right team usually beats hiring the single most optimal top talent — though for maybe 5% of roles where speed to market is critical, you do need proven experience. Because he interviews across every kind of expertise, he reduces it to three things: are you a curious problem solver who can articulate it, can you work humbly across people, and are you a good leader. His Uber Eats management team stayed largely intact from zero to $20B.

  • For ~5% of roles (e.g. researchers in a fast market) you need proven experience and relationships
  • For everyone else he interviews for three things: curious problem solver, works across people humbly, good leader
  • He composes a team like 'an organism of strengths' while minimizing conflicts
  • The Uber Eats management team was largely the same from nothing to $20B
  • A team that knows each other's strengths and can compensate beats hiring for who has 'seen this much scale'

I just believed that the team knowing each other's strengths and weaknesses and being able to compensate for each other was more important than the…

Jason Droege · 1:11:00
#hiring#teams#leadership#management