LLenny's Podcast
← All frameworks
Peak PerformanceEdwin Chen (Surge AI)

The Complementary Learning Modes Ladder

Mastery needs many kinds of practice — imitate a master, get preference feedback, get graded, then live it.

Difficulty
Moderate
Time to result
~ongoing to results
Steps
5
Confidence
83%

A model of how mastery is actually acquired, mapped from the progression of AI post-training methods onto the many ways humans learn. No single mode is enough and none makes the others obsolete — each teaches a different skill. Use it when designing a training program, learning a hard skill, or teaching someone (or something) to be genuinely good rather than just compliant.

Origin

Chen walks through the evolution of AI post-training — SFT, then RLHF, then rubrics and verifiers, now RL environments — using human learning analogies throughout, and argues you need a suite of methods mirroring the thousand different ways a great writer becomes great.

Core principles

  • 01There are a million ways humans learn; no single method is sufficient.
  • 02New learning modes complement rather than replace the older ones.
  • 03You develop taste by exposure to both masterpieces and terrible work.
  • 04Mastery comes from an endless cycle of practice and reflection.

How to run it

  1. 1

    Imitate a master (SFT)

    Start by copying what an expert does — supervised fine-tuning is analogous to mimicking a master and reproducing their moves.

  2. 2

    Produce many attempts and get preference feedback (RLHF)

    Write 55 different essays and have someone tell you which they like most — learning from comparative preference across many of your own attempts.

  3. 3

    Get graded with detailed feedback (rubrics and verifiers)

    Be evaluated against explicit criteria and receive detailed feedback on exactly where you went wrong.

  4. 4

    Practice in a realistic environment (RL environments)

    Get thrown into a messy simulation of the real task and learn by trying things end-to-end, seeing what works and reflecting.

    Pro tip This is closest to how humans actually learn — try stuff, figure out what's working, reflect.

  5. 5

    Keep all modes running

    Don't discard earlier modes as you add new ones — each remains a distinct skill the learner needs; run them as a complementary suite.

    Pro tip Develop taste by exposing yourself to both masterpieces and terrible examples, not just rules.

In the wild

Becoming a great writer

You don't become a great writer by memorizing grammar rules. You read great books, practice writing, get feedback from teachers and from readers who leave reviews, and develop taste by seeing both masterpieces and terrible writing — an endless cycle of practice and reflection.

A thousand different learning mechanisms combine into mastery, none replacing the others.

The post-training progression

AI went from SFT (mimic a master) to RLHF (pick the best of many attempts) to rubrics/verifiers (graded feedback) to RL environments (practice in a simulated world) — each a new complementary skill.

Chen stresses previous methods aren't obsolete; the new one is 'just a different form of learning.'

Common mistakes

Relying on rules alone

Memorizing grammar rules doesn't make a great writer — mastery needs practice, feedback, and exposure, not just explicit instruction.

Assuming the newest method replaces the rest

Each new learning mode complements the old ones; discarding earlier modes drops skills the learner still needs.

Is it for you?

Best for

Educators, coaches, and self-directed learners designing how to develop genuine mastery of a complex skill.

Not ideal for

Simple, rote skills where a single practice mode is entirely sufficient.

From the transcript

SFT is a lot by is a lot like mimicking a master and copying what they do

41:30

sometimes you learn by writing 55 different essays

41:30

rubrics and verifiers are like learning by being graded and getting detailed feedback on where you went wrong

42:00

So you learn through this endless cycle of practice and reflection.

43:30

From the episode

The 100-person AI lab that became Anthropic and Google's secret weapon

Edwin Chen (Surge AI)