LLenny's Podcast
← All frameworks
LeadershipHamel Husain & Shreya Shankar (creators of the #1 eval cours

The Benevolent Dictator Labeling Model

Appoint one trusted domain expert to own eval judgments instead of running it by committee.

Difficulty
Starter
Time to result
~days to results
Steps
3
Confidence
85%

A team-process decision rule for who makes the subjective quality calls during error analysis. Rather than convening a committee to reach consensus on what counts as a failure, you appoint a single person whose taste you trust — ideally the domain expert (often the product manager) — to be the 'benevolent dictator' doing the note-taking. The point is to keep the process cheap enough that it actually happens.

Origin

Hamel Husain coined the term 'benevolent dictator' for this context; Shreya Shankar confirms he came up with it. The underlying tension (consensus vs. speed in qualitative labeling) is a general research-ops problem.

Core principles

  • 01Committees make labeling so expensive the process dies — cut through the noise
  • 02The goal is progress and quick signal, not perfection or fairness
  • 03Judgments should be binary (good enough / not) rather than 1-5 scores
  • 04The labeler must have genuine domain context to judge product quality

How to run it

  1. 1

    Reject the committee instinct

    Resist the organizational urge to get everyone on board and involved in labeling. For most small-to-medium teams this is wholly unnecessary and stalls the work.

    Watch out Making the process a committee is the fastest way to make it too expensive to run at all — you'll lose out entirely.

  2. 2

    Appoint one domain expert as the dictator

    Pick a single person whose taste you trust and who has real domain context. For legal it's a lawyer, for mental health a clinician, and very often for a product it's the product manager. They own the open-coding judgments.

    Pro tip The domain expert is right because the AI conversation is the actual user experience — someone who understands the business (e.g. apartment leasing) can tell when a technically-fine answer is a product failure.

  3. 3

    Force binary judgments

    Have the dictator make yes/no calls, not graded scores. Binary keeps the process tractable and prevents endless deliberation over whether something is a 3 or a 4.

    Watch out Numeric rating scales slow the process to a crawl and produce metrics (3.2 vs 3.7) nobody can interpret or act on.

In the wild

Nurture Boss leasing judgments

For the apartment-leasing assistant, the benevolent dictator would be someone who understands the business of apartment leasing and has the context to judge whether a given AI response makes sense — deciding, for instance, that 'we don't have that available' without a human handoff is a product failure.

One trusted expert can drive the entire labeling effort, keeping error analysis cheap and repeatable instead of bottlenecked on group consensus.

Common mistakes

Demanding fairness over progress

It may feel unfair that one person is 'the dictator,' but optimizing for buy-in and fairness reintroduces the committee cost the model exists to avoid. The tradeoff is intentional: speed and signal now.

Is it for you?

Best for

Small-to-medium AI teams where labeling has stalled because too many stakeholders want a say in defining quality.

Not ideal for

Situations with no single trustworthy domain expert, or high-stakes domains where a single person's taste shouldn't be the sole arbiter of quality.

From the transcript

a lot of teams get bogged down in having a committee do this. And for a lot of situations that's wholly unnecessary

25:30

you can appoint one person whose taste that you trust.

26:00

it should be the person with domain expertise.

27:00

oftentimes it is the product manager.

27:30

From the episode

Why AI evals are the hottest new skill for product builders

Hamel Husain & Shreya Shankar (creators of the #1 eval cours