The Complementary Learning Modes Ladder
Mastery needs many kinds of practice — imitate a master, get preference feedback, get graded, then live it.
- Difficulty
- Moderate
- Time to result
- ~ongoing to results
- Steps
- 5
- Confidence
- 83%
A model of how mastery is actually acquired, mapped from the progression of AI post-training methods onto the many ways humans learn. No single mode is enough and none makes the others obsolete — each teaches a different skill. Use it when designing a training program, learning a hard skill, or teaching someone (or something) to be genuinely good rather than just compliant.
Origin
Chen walks through the evolution of AI post-training — SFT, then RLHF, then rubrics and verifiers, now RL environments — using human learning analogies throughout, and argues you need a suite of methods mirroring the thousand different ways a great writer becomes great.
Core principles
- 01There are a million ways humans learn; no single method is sufficient.
- 02New learning modes complement rather than replace the older ones.
- 03You develop taste by exposure to both masterpieces and terrible work.
- 04Mastery comes from an endless cycle of practice and reflection.
How to run it
- 1
Imitate a master (SFT)
Start by copying what an expert does — supervised fine-tuning is analogous to mimicking a master and reproducing their moves.
- 2
Produce many attempts and get preference feedback (RLHF)
Write 55 different essays and have someone tell you which they like most — learning from comparative preference across many of your own attempts.
- 3
Get graded with detailed feedback (rubrics and verifiers)
Be evaluated against explicit criteria and receive detailed feedback on exactly where you went wrong.
- 4
Practice in a realistic environment (RL environments)
Get thrown into a messy simulation of the real task and learn by trying things end-to-end, seeing what works and reflecting.
Pro tip This is closest to how humans actually learn — try stuff, figure out what's working, reflect.
- 5
Keep all modes running
Don't discard earlier modes as you add new ones — each remains a distinct skill the learner needs; run them as a complementary suite.
Pro tip Develop taste by exposing yourself to both masterpieces and terrible examples, not just rules.
In the wild
You don't become a great writer by memorizing grammar rules. You read great books, practice writing, get feedback from teachers and from readers who leave reviews, and develop taste by seeing both masterpieces and terrible writing — an endless cycle of practice and reflection.
→ A thousand different learning mechanisms combine into mastery, none replacing the others.
AI went from SFT (mimic a master) to RLHF (pick the best of many attempts) to rubrics/verifiers (graded feedback) to RL environments (practice in a simulated world) — each a new complementary skill.
→ Chen stresses previous methods aren't obsolete; the new one is 'just a different form of learning.'
Common mistakes
Relying on rules alone
Memorizing grammar rules doesn't make a great writer — mastery needs practice, feedback, and exposure, not just explicit instruction.
Assuming the newest method replaces the rest
Each new learning mode complements the old ones; discarding earlier modes drops skills the learner still needs.
Is it for you?
Best for
Educators, coaches, and self-directed learners designing how to develop genuine mastery of a complex skill.
Not ideal for
Simple, rote skills where a single practice mode is entirely sufficient.
From the transcript
“SFT is a lot by is a lot like mimicking a master and copying what they do”
“sometimes you learn by writing 55 different essays”
“rubrics and verifiers are like learning by being graded and getting detailed feedback on where you went wrong”
“So you learn through this endless cycle of practice and reflection.”
From the episode
The 100-person AI lab that became Anthropic and Google's secret weapon
Edwin Chen (Surge AI)