LLenny's Podcast
← All episodes
Brendan Foody (CEO of Mercor)18 September 2025

Why experts writing AI evals is creating the fastest-growing companies in history

7Frameworks
14Insights

Episode overview

Lenny Rachitsky interviews Mercor co-founder and CEO Brendan Foody about why expert-written evaluations have become a central bottleneck in improving AI models. Foody explains how Mercor’s marketplace recruits skilled professionals to define success criteria and post-training data, helping the company grow from roughly $1 million to $400 million in annualized revenue in 16 months. They also discuss AI’s effects on work, the value of highly elastic industries, Mercor’s operating principles, and why Foody expects progress toward superintelligence to take longer than some lab leaders predict.

Key ideas

  • When the model is the product, an evaluation functions like its product requirements document and can also become sales collateral.
  • AI labs increasingly need skilled professionals—such as lawyers, doctors, engineers, and creatives—to define rubrics and verifiers for model capabilities.
  • The market for expert training work persists as long as humans can perform economically valuable tasks that models cannot.
  • Reinforcement learning from AI feedback replaces repeated human preference judgments with human-defined success criteria that can scale more efficiently.
  • Workers and companies should learn to use AI to expand output, especially in elastic fields such as software and product development.
  • Mercor found product-market fit by following strong market pull, obsessing over customers, and treating experts better than legacy crowdsourcing platforms.
  • Mercor’s cultural principles are a can-do attitude, high standards, and intensity measured through output rather than mandated working hours.
  • Foody believes most knowledge-work tasks may be automated within a decade, but that superintelligence is likely farther away than three years.

Transcript available · source text is retained privately and is not published

Frameworks in this episode

People & resources mentioned

Attributed to the moment in the episode. Timestamps are approximate.

People · 12

  • Brendan FoodyMentionsToday, my guest is Brendan Foody, CEO and co-founder of Merkor.

    Today, my guest is Brendan Foody, CEO and co-founder of Merkor.

  • Lenny RachitskyMentionstoday we've got another very special compilation episode something I've been pulling on more and more with the podcast and the newsletter

    Thank you so much for having me, Lenny.

  • Sarah GuoMentionsWhat a legendary panel we assembled there with Sarah Guo moderating.

    I saw you talking about this on the No Priors podcast with Sarah and Ilad

  • Elad GilMentionsI really love elad Gill's High grth handbook

    I saw you talking about this on the No Priors podcast with Sarah and Ilad

  • Andrej KarpathyMentionsI remember Karpathy tweeted that he just like has never seen a model like this.

    one thing like Andre Karpathy tweeted about a bunch

  • Marc AndreessenMentionsMark Andre had this visual on some podcasts have just imagined a 100,000 drones just coming out of China just at us.

    Mark and tweeted about this recently that software is the most elastic industry of all

  • Elon MuskMentionsWhether you like him or not, Elon, this is similar to his um if you watch him, he has sets these really ambitious goals

    This reminds me of Elon's whole thing with Neurolink

  • Greg BrockmanMentionsthere's everybody has uh and had and continues to have like such deep trust in in Sam and Greg and our leadership team

    I think Greg Greg Brockman tweeted this once. Evals are all you need.

  • Elon MuskMentionsWhether you like him or not, Elon, this is similar to his um if you watch him, he has sets these really ambitious goals

    we met the entire XAI co-founding team except for Elon

  • Sundeep JainMentionsHe was previously the chief product officer and chief technology officer at Uber

    We just hired uh or partnered with Sundep Jane who joined us as president.

  • David SacksMentionsceo of paypal and then the founder of yammer

    David Sax tweeted this interesting point that like the situation we're in now is almost the best case scenario

  • Eric AntonowMentionsHe's this creative uh product person that's kind of under the radar now. He's at Facebook for a long time.

    there's this uh guy, Eric Antinau, who recommended by a lot of people to get him on this podcast.

Resources · 40

  • MercorCoinedcompany · Brendan Foody and co-founders

    Today, my guest is Brendan Foody, CEO and co-founder of Merkor.

  • Lenny's NewsletterCoinednewsletter · Lenny Rachitsky

    if you become an annual subscriber of my newsletter, you get 15 incredible products for free for 1 year

  • ZanzibarCoinedsoftware · Google

    Zanzibar, which was originally designed for Google to power Google Docs and YouTube.

  • WorkOSRecommendscompany · WorkOS, Inc.

    you should consider work OS.

  • WarrantMentionssoftware · Warrant

    Work OS also recently acquired Warrant, the fine grain authorization service.

  • Jira Product DiscoveryRecommendssoftware · Atlassian

    Rediscover what's possible with Jira product discovery. Try it for free at atlassian.com/lenny.

  • SWE-benchMentionsother

    how fast we're even saturating Sweetbench once we focus on it.

  • No PriorsMentionspodcast · Sarah Guo

    I saw you talking about this on the No Priors podcast with Sarah and Ilad

  • GPQAMentionspaper

    everyone's been pointing to these academic evals of PhD level reasoning with GPQA, humanities last exam or Olympiad math.

  • Humanity's Last ExamMentionsother

    everyone's been pointing to these academic evals of PhD level reasoning with GPQA, humanities last exam or Olympiad math.

  • HandshakeMentionscompany · Garrett Lord and co-founders

    I've had the CEO of Handshake on the podcast.

  • OpenAIMentionscompany

    but we met OpenAI and we saw that there was this enormous transition in the human data market

  • AnthropicMentionscompany

    That's what they've done at Anthropic is move towards AIdriven reinforcement learning.

  • Claude CodeUsessoftware · Anthropic

    use cloud code use whatever tool cursor and whatever tools are available to build a website

  • CursorUsessoftware · Anysphere

    use cloud code use whatever tool cursor and whatever tools are available to build a website

  • McKinsey & CompanyMentionscompany

    imagine if we could have 10 times as many Mckenzie consultants.

  • NeuralinkMentionscompany · Elon Musk and co-founders

    This reminds me of Elon's whole thing with Neurolink

  • EnterpretMentionssoftware

    Interpret is a customer intelligence platform used by leading CX and product orgs

  • Surge AIMentionscompany · Edwin Chen

    I think like scale and surge were the primary companies that pioneered that industry.

  • Scale AIMentionscompany · Alex Wang and Lucy Guo

    I think like scale and surge were the primary companies that pioneered that industry.

  • AlphaSightsMentionscompany

    There's always been these companies, Alphasites and uh GLG

  • GLGMentionscompany

    There's always been these companies, Alphasites and uh GLG

  • Harvard LampoonMentionscompany

    Like we hired all the people from the Harvard Lampoon a couple months ago

  • Goldman SachsMentionscompany

    the, you know, Goldman, the Mckenzie analysts, the, uh, Fang software engineers.

  • xAIMentionscompany · Elon Musk

    one of our customers introduced us to the co-founders of XAI

  • TeslaMentionscompany · Tesla, Inc.

    So, they had us in uh two days later to the Tesla office

  • General CatalystMentionscompany

    we raised our seed round in September from general catalyst

  • BenchmarkMentionscompany

    when we were talking to Benchmark uh before they let our series A

  • UberMentionscompany · Uber

    He was previously the chief product officer and chief technology officer at Uber

  • Doughnut DynastyCoinedcompany · Brendan Foody

    when I was in eighth grade, I started Doughnut Dynasty

  • ChatGPTUsessoftware · OpenAI

    I like chatt voice mode a lot.

  • ParrotGPTCoinedproduct · Eric Antonow

    He built this project called Parrot GPT

  • SuitsRecommendstv show · Aaron Korsh

    My favorite TV show of all time is Suits.

  • High Output ManagementRecommendsbook · Andy Grove

    I would say in order high output management is a phenomenal book on running companies.

  • Zero to OneRecommendsbook · Peter Thiel and Blake Masters

    Second is 0ero to1 uh which of course is a classic.

  • Shoe DogRecommendsbook · Phil Knight

    And then third is shoe dog where I just find it to be a really inspirational story.

  • OppenheimerRecommendsfilm · Christopher Nolan

    I really liked Oenheimer.

  • CodexUsessoftware · OpenAI

    I love using codecs, like the new version.

  • MercorMentionscompany · Brendan Foody and co-founders

    please go to mercur.com and we would love to work with you.

  • Lenny's PodcastCoinedpodcast · Lenny Rachitsky

    You can find all past episodes or learn more about the show at lennispodcast.com.

Spot an error or want something removed? Request a correction or removal.

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Hot Take· 2

Hot Take24:00

The Most Successful People Won't Have a Specific Skill, They'll Be Good With AI

Brendan's advice for future-proofing isn't a particular discipline but fluency with AI itself, using the tools to get better at whatever you already do. He points to how Mercor now runs interviews telling candidates to use ChatGPT, Claude Code, Cursor, and whatever tools they want to build a website in an hour, rather than fighting tool use the way schools once fought calculators.

  • The differentiating skill is being good with AI, not any single discipline.
  • Don't fight people using models; it's like fighting the calculator.
  • Mercor's interviews tell candidates to use any AI tools to build a product in an hour.
  • Leverage the technology to do much more in whatever vertical you operate in.

It's not like a specific skill, but it's being good with AI, using AI to become more become better at what you're already doing.

Brendan Foody · 24:00
#ai#skills#hiring#assessment
Hot Take58:00

Super Intelligence Isn't Three Years Away: It's a Longer Road Paved With Evals

Brendan pushes back on lab executives predicting super intelligence within three years, arguing the truth is a longer road. He still expects models to automate a majority of knowledge-work tasks within 10 years, but says that progress comes from data-efficient post-training datasets and evals, not from 10x more pre-training data.

  • Predictions of super intelligence in three years are too optimistic.
  • Models will likely automate a majority of knowledge work within 10 years.
  • Progress comes from thoughtful, data-efficient post-training, not 10x pre-training data.
  • That long road is paved with the evals that make new capabilities possible.

I know there's been some executives at big labs that say we'll have super intelligence in three years, but I think the truth is that…

Brendan Foody · 58:00

that long road is paved with all of the evals that help to make those capabilities possible.

Brendan Foody · 58:30
#agi#super-intelligence#scaling-laws#predictions

Explainer· 5

Explainer06:30

The Era of Evals: If the Model Is the Product, the Eval Is the PRD

Brendan explains why 'evals' have become the central bottleneck in AI. If the model is the product, the eval is the product requirement document that tells researchers what to build. Because reinforcement learning has gotten so effective, once labs have an eval they can rapidly 'hill climb' it, the way they saturated Olympiad math and SWE-bench once they focused on them.

  • An eval is a systematic way to measure what success looks like for a model.
  • Once labs have an eval, RL lets them hill-climb capability quickly.
  • The barrier to automating any workflow is figuring out how to measure success.
  • Mercor works with six of the magnificent seven and all top five AI labs.

If the model is the product, then the eval is the product requirement document.

Brendan Foody · 06:30

once they have an EVL, they can hill climb it.

Brendan Foody · 06:30
#ai#evals#reinforcement-learning#model-training
Explainer13:00

What Experts Actually Do: A Lawyer Writing a Rubric for a Contract

Brendan makes the abstract concrete: the market is bound by the number of things humans can do that models can't. To improve a model that botches a contract redline, a lawyer creates a rubric, much like a professor grading a deliverable, that scores whether the model hit each key point. That rubric becomes both the measure of progress and the training signal to reward the desired capabilities.

  • The market is bounded by the gap between what humans can do and what models can't.
  • An expert writes a rubric that scores a model's output point by point.
  • The rubric doubles as the eval and as training data to reinforce capabilities.
  • This mirrors how a professor creates grading criteria for a deliverable.

Effectively, the market is bound by the amount of things where humans can do something that models can't.

Brendan Foody · 13:00

have a lawyer create a rubric similar to how a professor might create a rubric

Brendan Foody · 13:30
#evals#experts#rubrics#ai-training
Explainer15:30

The Shift From Human Feedback to Reinforcement Learning From AI Feedback

Brendan walks through the evolution of training data: from supervised fine-tuning (input/output pairs) to RLHF (humans pick the best of several outputs) to the current move toward reinforcement learning from AI feedback (RLAIF). In RLAIF the human defines success criteria (a unit test, a rubric) and the model is rewarded against it, which is far more scalable and data-efficient.

  • Supervised fine-tuning is input/output data; RLHF has humans rank generated examples.
  • The field is moving toward reinforcement learning from AI feedback (RLAIF).
  • In RLAIF, humans define success criteria (unit tests, rubrics) rather than label every example.
  • RLAIF is more scalable and data-efficient for both evaluating and improving models.

What everyone is generally moving towards is reinforcement learning from AI feedback instead of human feedback.

Brendan Foody · 15:30
#rlhf#rlaif#training-data#fine-tuning
Explainer30:30

How the Model Learned to Read Your X-Ray: Post-Training, Not the Internet

Prompted by a story of ChatGPT reading a friend's X-ray, Brendan explains the split: pre-training loads broad knowledge from the world, while post-training and reinforcement learning teach the model which pieces are accurate and how to reason. Behind that X-ray capability were radiologists who built the post-training dataset defining the right diagnoses and the rewards and penalties around them.

  • Pre-training loads knowledge; post-training and RL teach reasoning and prioritization.
  • Experts don't feed the model new data; they refine what it already knows.
  • Radiologists built the post-training dataset behind medical-image capabilities.
  • The quality of those experts drives the quality of the model's recommendations.

behind that there would have been radiologists that uh worked on the post-training data set to create some stasis point for here's the diagnosis and…

Brendan Foody · 30:30
#pre-training#post-training#experts#reasoning
Explainer37:00

What Experts Get Paid: $95/hr Median, Up to $500/hr

Brendan shares the economics that make this meaningful even for FAANG engineers: the median pay rate in Mercor's marketplace is $95 an hour, flexing up to around $500 an hour depending on depth of expertise. That contrasts sharply with legacy crowdsourcing companies that averaged around $30 an hour for undergrads, because labs want Goldman, McKinsey, and FAANG-caliber talent.

  • Median pay is $95/hour, flexing up to ~$500/hour based on expertise.
  • Legacy crowdsourcing paid roughly $30/hour on average.
  • Labs want Goldman/McKinsey analysts and FAANG engineers, not undergrads.
  • Work is often part-time (e.g. an underemployed FAANG engineer with spare hours).

our median pay rate in the marketplace is $95 an hour, but it can flex up well up into like $500 an hour um based…

Brendan Foody · 37:00
#pay#marketplace#experts#economics

Story· 4

Story11:00

1 to $400M Revenue in 16 Months: The Fastest Ascent in History

Brendan recounts how Mercor started as an international hiring company, bootstrapped to $1M run rate before he dropped out of college, then pivoted after meeting OpenAI and spotting a shift in the human-data market from crowdsourcing low-skill labor to sourcing and vetting elite professionals. From there they grew from $1M to $400M in revenue run rate in 16 months.

  • The human-data market shifted from crowdsourcing low/medium-skill workers to sourcing and vetting top experts.
  • Mercor bootstrapped to $1M run rate before the founders dropped out of college.
  • They grew from $1M to $400M revenue run rate in 16 months.
  • Brendan met his co-founders at age 14 and started the company at 19.

We grew from 1 to 400 million in revenue run rate in uh 16 months.

Brendan Foody · 11:00

So fastest ascent uh in history uh which is an exciting statistic we're we're very proud of.

Brendan Foody · 11:30
#startups#growth#mercor#founding-story
Story34:00

Hiring the Harvard Lampoon to Make Models Funnier

Beyond hard professional domains, labs still care about creative capability. Brendan shares that Mercor hired the entire Harvard Lampoon comedy club a couple of months earlier to help make models funnier, and brings on Emmy-award-winning screenwriters and other creative talent across the board.

  • Labs want both professional/economic and creative capabilities.
  • Mercor hired the whole Harvard Lampoon comedy club to make models funnier.
  • They also hire Emmy-award-winning screenwriters for creative work.
  • A customer request can be turned around with experts within 24 hours.

Like we hired all the people from the Harvard Lampoon a couple months ago, their comedy club to help with making models funnier.

Brendan Foody · 34:30
#hiring#creative#experts#mercor
Story53:30

Doughnut Dynasty: The 8th-Grade Business That Taught Him You Can Just Do Things

In eighth grade Brendan noticed Safeway sold donuts for $5 a dozen, so he bought them and resold them at school for $2 each, scaled by paying his mom $20 to drive him, undercut a competitor to $1 to run them out of business, and paid friends in donuts. The takeaway: you can just do things, and the real barrier to more companies is initiative, not ideas.

  • Bought Safeway donuts at $5/dozen and resold them at $2 each with strong margins.
  • Paid his mom $20 to drive the minivan to buy 10 dozen at a time.
  • Dropped prices to $1 to run a higher-cost competitor out of business.
  • The barrier to building companies is initiative, not a shortage of ideas.

you can just do things. Like so many people have ideas, but the barrier to more companies being built I think is just initiative and…

Brendan Foody · 55:00
#entrepreneurship#story#initiative#childhood-business
Story1:05:00

Being Dyslexic Made Him See Markets Differently, and Manage to Strengths

Brendan openly shares that he's dyslexic. While it makes reading a thousand emails a day or every document hard, he believes it helps him think differently, be more creative, and see market shifts others miss. It shaped a management philosophy of leveraging people's strengths rather than trying to fix their weaknesses.

  • Brendan is dyslexic and doesn't hide it from colleagues.
  • It makes reading heavy volumes of email and documents difficult.
  • He credits it with more creative thinking and spotting market changes early.
  • It informs his focus on leveraging strengths over fixing weaknesses.

But on the other hand, I feel like it helps me to think a little bit differently, to be more creative, and perhaps see the…

Brendan Foody · 1:05:00

we focus much more on how we can leverage people's strengths rather than helping to improve weaknesses.

Brendan Foody · 1:05:30
#dyslexia#management#strengths#leadership

Takeaway· 3

Takeaway20:30

Which Jobs Survive AI: Bet on Industries With Elastic Demand

Asked which skills are worth investing in, Brendan argues the winning categories are those with elastic demand, where making people more productive increases demand rather than shrinking it. Accounting is inelastic (the world doesn't need 100x more), but software is the most elastic industry of all, so learning to code, product management, and operations still pay off.

  • Bet on industries where 10x productivity increases demand rather than reducing it.
  • Accounting is inelastic; the world doesn't need 100x more of it.
  • Software is the most elastic industry; more productivity means far more gets built.
  • Product managers who can now do far more are extremely well positioned.

software is the most elastic industry of all where when we increase productivity, there's so much more that will be built.

Brendan Foody · 23:00
#careers#jobs#ai#software
Takeaway35:30

The Top 10% of Experts Drive the Majority of Model Improvement

Brendan describes a power-law dynamic in the work: in any batch of 100 people they hire, the top 10% drive the majority of the model improvement, just like the top 10% of a company drives most of its impact. Building proprietary advantages in identifying and matching those top-10% people is what makes Mercor hard to compete against.

  • In a batch of 100 hires, the top 10% drive most of the model improvement.
  • This mirrors how the top 10% of a company drives most of its impact.
  • The moat is proprietary ability to identify and match those top performers.
  • This ties back to the founding thesis of finding extraordinary people.

in a set of a 100 people that we hire oftent times the top 10% of people will drive majority of the model improvement.

Brendan Foody · 35:30
#talent#power-law#moat#hiring
Takeaway44:00

Stop Forcing Product-Market Fit: Find the Customer Who's Surprisingly Easy to Sell

Brendan's advice for founders is that he wasted time trying to force product-market fit. Instead of pushing a hard-to-sell product, you should find the customer who is surprisingly easy to sell into and can grow with you. The balance is being stubborn about your thesis for how the world changes while staying open-minded about the exact form it takes.

  • Trying to force product-market fit is a common founder mistake.
  • Look for the customer who is surprisingly easy to sell into.
  • If the marginal customer is very hard to sell, you can't build a huge business.
  • Be stubborn on your thesis but open-minded on how it plays out.

What you actually need to find is the customer that's surprisingly easy to sell into where you're going to be able to grow with them.

Brendan Foody · 44:00

it's some combination of being stubborn with respect to your thesis around how the world will change, but also very open-minded with respect to exactly…

Brendan Foody · 44:30
#founders#product-market-fit#sales#advice