LLenny's Podcast
← All episodes
Brendan Foody (CEO of Mercor)18 September 2025

Why experts writing AI evals is creating the fastest-growing companies in history

7Frameworks
14Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Hot Take· 2

Hot Take24:00

The Most Successful People Won't Have a Specific Skill, They'll Be Good With AI

Brendan's advice for future-proofing isn't a particular discipline but fluency with AI itself, using the tools to get better at whatever you already do. He points to how Mercor now runs interviews telling candidates to use ChatGPT, Claude Code, Cursor, and whatever tools they want to build a website in an hour, rather than fighting tool use the way schools once fought calculators.

  • The differentiating skill is being good with AI, not any single discipline.
  • Don't fight people using models; it's like fighting the calculator.
  • Mercor's interviews tell candidates to use any AI tools to build a product in an hour.
  • Leverage the technology to do much more in whatever vertical you operate in.

It's not like a specific skill, but it's being good with AI, using AI to become more become better at what you're already doing.

Brendan Foody · 24:00
#ai#skills#hiring#assessment
Hot Take58:00

Super Intelligence Isn't Three Years Away: It's a Longer Road Paved With Evals

Brendan pushes back on lab executives predicting super intelligence within three years, arguing the truth is a longer road. He still expects models to automate a majority of knowledge-work tasks within 10 years, but says that progress comes from data-efficient post-training datasets and evals, not from 10x more pre-training data.

  • Predictions of super intelligence in three years are too optimistic.
  • Models will likely automate a majority of knowledge work within 10 years.
  • Progress comes from thoughtful, data-efficient post-training, not 10x pre-training data.
  • That long road is paved with the evals that make new capabilities possible.

I know there's been some executives at big labs that say we'll have super intelligence in three years, but I think the truth is that…

Brendan Foody · 58:00

that long road is paved with all of the evals that help to make those capabilities possible.

Brendan Foody · 58:30
#agi#super-intelligence#scaling-laws#predictions

Explainer· 5

Explainer06:30

The Era of Evals: If the Model Is the Product, the Eval Is the PRD

Brendan explains why 'evals' have become the central bottleneck in AI. If the model is the product, the eval is the product requirement document that tells researchers what to build. Because reinforcement learning has gotten so effective, once labs have an eval they can rapidly 'hill climb' it, the way they saturated Olympiad math and SWE-bench once they focused on them.

  • An eval is a systematic way to measure what success looks like for a model.
  • Once labs have an eval, RL lets them hill-climb capability quickly.
  • The barrier to automating any workflow is figuring out how to measure success.
  • Mercor works with six of the magnificent seven and all top five AI labs.

If the model is the product, then the eval is the product requirement document.

Brendan Foody · 06:30

once they have an EVL, they can hill climb it.

Brendan Foody · 06:30
#ai#evals#reinforcement-learning#model-training
Explainer13:00

What Experts Actually Do: A Lawyer Writing a Rubric for a Contract

Brendan makes the abstract concrete: the market is bound by the number of things humans can do that models can't. To improve a model that botches a contract redline, a lawyer creates a rubric, much like a professor grading a deliverable, that scores whether the model hit each key point. That rubric becomes both the measure of progress and the training signal to reward the desired capabilities.

  • The market is bounded by the gap between what humans can do and what models can't.
  • An expert writes a rubric that scores a model's output point by point.
  • The rubric doubles as the eval and as training data to reinforce capabilities.
  • This mirrors how a professor creates grading criteria for a deliverable.

Effectively, the market is bound by the amount of things where humans can do something that models can't.

Brendan Foody · 13:00

have a lawyer create a rubric similar to how a professor might create a rubric

Brendan Foody · 13:30
#evals#experts#rubrics#ai-training
Explainer15:30

The Shift From Human Feedback to Reinforcement Learning From AI Feedback

Brendan walks through the evolution of training data: from supervised fine-tuning (input/output pairs) to RLHF (humans pick the best of several outputs) to the current move toward reinforcement learning from AI feedback (RLAIF). In RLAIF the human defines success criteria (a unit test, a rubric) and the model is rewarded against it, which is far more scalable and data-efficient.

  • Supervised fine-tuning is input/output data; RLHF has humans rank generated examples.
  • The field is moving toward reinforcement learning from AI feedback (RLAIF).
  • In RLAIF, humans define success criteria (unit tests, rubrics) rather than label every example.
  • RLAIF is more scalable and data-efficient for both evaluating and improving models.

What everyone is generally moving towards is reinforcement learning from AI feedback instead of human feedback.

Brendan Foody · 15:30
#rlhf#rlaif#training-data#fine-tuning
Explainer30:30

How the Model Learned to Read Your X-Ray: Post-Training, Not the Internet

Prompted by a story of ChatGPT reading a friend's X-ray, Brendan explains the split: pre-training loads broad knowledge from the world, while post-training and reinforcement learning teach the model which pieces are accurate and how to reason. Behind that X-ray capability were radiologists who built the post-training dataset defining the right diagnoses and the rewards and penalties around them.

  • Pre-training loads knowledge; post-training and RL teach reasoning and prioritization.
  • Experts don't feed the model new data; they refine what it already knows.
  • Radiologists built the post-training dataset behind medical-image capabilities.
  • The quality of those experts drives the quality of the model's recommendations.

behind that there would have been radiologists that uh worked on the post-training data set to create some stasis point for here's the diagnosis and…

Brendan Foody · 30:30
#pre-training#post-training#experts#reasoning
Explainer37:00

What Experts Get Paid: $95/hr Median, Up to $500/hr

Brendan shares the economics that make this meaningful even for FAANG engineers: the median pay rate in Mercor's marketplace is $95 an hour, flexing up to around $500 an hour depending on depth of expertise. That contrasts sharply with legacy crowdsourcing companies that averaged around $30 an hour for undergrads, because labs want Goldman, McKinsey, and FAANG-caliber talent.

  • Median pay is $95/hour, flexing up to ~$500/hour based on expertise.
  • Legacy crowdsourcing paid roughly $30/hour on average.
  • Labs want Goldman/McKinsey analysts and FAANG engineers, not undergrads.
  • Work is often part-time (e.g. an underemployed FAANG engineer with spare hours).

our median pay rate in the marketplace is $95 an hour, but it can flex up well up into like $500 an hour um based…

Brendan Foody · 37:00
#pay#marketplace#experts#economics

Story· 4

Story11:00

1 to $400M Revenue in 16 Months: The Fastest Ascent in History

Brendan recounts how Mercor started as an international hiring company, bootstrapped to $1M run rate before he dropped out of college, then pivoted after meeting OpenAI and spotting a shift in the human-data market from crowdsourcing low-skill labor to sourcing and vetting elite professionals. From there they grew from $1M to $400M in revenue run rate in 16 months.

  • The human-data market shifted from crowdsourcing low/medium-skill workers to sourcing and vetting top experts.
  • Mercor bootstrapped to $1M run rate before the founders dropped out of college.
  • They grew from $1M to $400M revenue run rate in 16 months.
  • Brendan met his co-founders at age 14 and started the company at 19.

We grew from 1 to 400 million in revenue run rate in uh 16 months.

Brendan Foody · 11:00

So fastest ascent uh in history uh which is an exciting statistic we're we're very proud of.

Brendan Foody · 11:30
#startups#growth#mercor#founding-story
Story34:00

Hiring the Harvard Lampoon to Make Models Funnier

Beyond hard professional domains, labs still care about creative capability. Brendan shares that Mercor hired the entire Harvard Lampoon comedy club a couple of months earlier to help make models funnier, and brings on Emmy-award-winning screenwriters and other creative talent across the board.

  • Labs want both professional/economic and creative capabilities.
  • Mercor hired the whole Harvard Lampoon comedy club to make models funnier.
  • They also hire Emmy-award-winning screenwriters for creative work.
  • A customer request can be turned around with experts within 24 hours.

Like we hired all the people from the Harvard Lampoon a couple months ago, their comedy club to help with making models funnier.

Brendan Foody · 34:30
#hiring#creative#experts#mercor
Story53:30

Doughnut Dynasty: The 8th-Grade Business That Taught Him You Can Just Do Things

In eighth grade Brendan noticed Safeway sold donuts for $5 a dozen, so he bought them and resold them at school for $2 each, scaled by paying his mom $20 to drive him, undercut a competitor to $1 to run them out of business, and paid friends in donuts. The takeaway: you can just do things, and the real barrier to more companies is initiative, not ideas.

  • Bought Safeway donuts at $5/dozen and resold them at $2 each with strong margins.
  • Paid his mom $20 to drive the minivan to buy 10 dozen at a time.
  • Dropped prices to $1 to run a higher-cost competitor out of business.
  • The barrier to building companies is initiative, not a shortage of ideas.

you can just do things. Like so many people have ideas, but the barrier to more companies being built I think is just initiative and…

Brendan Foody · 55:00
#entrepreneurship#story#initiative#childhood-business
Story1:05:00

Being Dyslexic Made Him See Markets Differently, and Manage to Strengths

Brendan openly shares that he's dyslexic. While it makes reading a thousand emails a day or every document hard, he believes it helps him think differently, be more creative, and see market shifts others miss. It shaped a management philosophy of leveraging people's strengths rather than trying to fix their weaknesses.

  • Brendan is dyslexic and doesn't hide it from colleagues.
  • It makes reading heavy volumes of email and documents difficult.
  • He credits it with more creative thinking and spotting market changes early.
  • It informs his focus on leveraging strengths over fixing weaknesses.

But on the other hand, I feel like it helps me to think a little bit differently, to be more creative, and perhaps see the…

Brendan Foody · 1:05:00

we focus much more on how we can leverage people's strengths rather than helping to improve weaknesses.

Brendan Foody · 1:05:30
#dyslexia#management#strengths#leadership

Takeaway· 3

Takeaway20:30

Which Jobs Survive AI: Bet on Industries With Elastic Demand

Asked which skills are worth investing in, Brendan argues the winning categories are those with elastic demand, where making people more productive increases demand rather than shrinking it. Accounting is inelastic (the world doesn't need 100x more), but software is the most elastic industry of all, so learning to code, product management, and operations still pay off.

  • Bet on industries where 10x productivity increases demand rather than reducing it.
  • Accounting is inelastic; the world doesn't need 100x more of it.
  • Software is the most elastic industry; more productivity means far more gets built.
  • Product managers who can now do far more are extremely well positioned.

software is the most elastic industry of all where when we increase productivity, there's so much more that will be built.

Brendan Foody · 23:00
#careers#jobs#ai#software
Takeaway35:30

The Top 10% of Experts Drive the Majority of Model Improvement

Brendan describes a power-law dynamic in the work: in any batch of 100 people they hire, the top 10% drive the majority of the model improvement, just like the top 10% of a company drives most of its impact. Building proprietary advantages in identifying and matching those top-10% people is what makes Mercor hard to compete against.

  • In a batch of 100 hires, the top 10% drive most of the model improvement.
  • This mirrors how the top 10% of a company drives most of its impact.
  • The moat is proprietary ability to identify and match those top performers.
  • This ties back to the founding thesis of finding extraordinary people.

in a set of a 100 people that we hire oftent times the top 10% of people will drive majority of the model improvement.

Brendan Foody · 35:30
#talent#power-law#moat#hiring
Takeaway44:00

Stop Forcing Product-Market Fit: Find the Customer Who's Surprisingly Easy to Sell

Brendan's advice for founders is that he wasted time trying to force product-market fit. Instead of pushing a hard-to-sell product, you should find the customer who is surprisingly easy to sell into and can grow with you. The balance is being stubborn about your thesis for how the world changes while staying open-minded about the exact form it takes.

  • Trying to force product-market fit is a common founder mistake.
  • Look for the customer who is surprisingly easy to sell into.
  • If the marginal customer is very hard to sell, you can't build a huge business.
  • Be stubborn on your thesis but open-minded on how it plays out.

What you actually need to find is the customer that's surprisingly easy to sell into where you're going to be able to grow with them.

Brendan Foody · 44:00

it's some combination of being stubborn with respect to your thesis around how the world will change, but also very open-minded with respect to exactly…

Brendan Foody · 44:30
#founders#product-market-fit#sales#advice