“Hamill and Shrea have played a major role in shifting evals from being an obscure mysterious subject”
Hamel Husain
Hamill and Shrea have played a major role in shifting evals from being an obscure mysterious subject
1 episode · 2 sourced appearances
Appearances and evidence
Short attributed excerpts only. Timestamps are approximate.
“They teach the definitive online course on evals, which happens to be the number one course on Maven.”
Frameworks in these episodes
Episode-linked context, not an unsupported claim of authorship.
Error Analysis: Open Coding to Axial Coding to Count
Turn messy LLM logs into a prioritized list of failures before you write a single test.
Eval Triage: Decide What Actually Deserves an Eval
After counting failures, route each one to a prompt fix, a code check, or an LLM judge — not all three.
The Aligned Binary LLM-as-Judge
Build a one-failure, pass/fail judge and align it to a human with a confusion matrix before you trust it.
The Benevolent Dictator Labeling Model
Appoint one trusted domain expert to own eval judgments instead of running it by committee.
Resources in their episodes
Related guests
Spot an error or want this profile removed? Request a correction or removal.