Evaluation and Measurement
AIE-1024 credits · 66h required · 1h optional · after AIE-101
You have shipped four LLM products and cannot currently answer, with a number, whether any of them is any good. This course fixes that, using those products as the laboratory. Nothing here is a toy dataset.
By the end you can
- Build an eval set from production failures rather than from imagination
- Name and measure the known pathologies of LLM judges — position, verbosity, self-enhancement bias
- Calibrate an automated judge against your own labels and report the agreement, not just the score
- Decide whether a difference between two systems is real, with an interval rather than a vibe
- Wire evals into CI so a regression fails a build
0 of 47 required items complete
0m of 66h 10m
M1 · What an eval is, and what it is not
0/6 · 6h 40mlecture · 1h 40m · tier 0 self-marked
Why AI evals are the hottest new skill for product builders — Lenny's Podcast
reading · 1h · tier 0 self-marked
reading · 45m · tier 0 self-marked
reading · 1h 30m · tier 0 self-marked
reading · 1h · tier 0 self-marked
assignment · 45m · tier 2 panel-assessed
M2 · Error analysis and the first eval set
0/6 · 9hlecture · 1h 30m · tier 0 self-marked
How to Build and Evaluate AI systems in the Age of LLMs — DataTalksClub ⬛
assignment · 3h · tier 2 panel-assessed
assignment · 3h · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
assignment · 30m · tier 2 panel-assessed
reading · 45m · tier 0 self-marked
M3 · LLM-as-judge and its pathologies
0/10 · 14hlecture · 1h · tier 0 self-marked
LLMOps (LLM Bootcamp) — The Full Stack
reading · 2h · tier 0 self-marked
reading · 1h · tier 0 self-marked
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
reading · 1h · tier 0 self-marked
reading · 1h · tier 0 self-marked
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
reading · 1h 15m · tier 0 self-marked
assignment · 2h 30m · tier 1 machine-verified
assignment · 3h · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
reading · 1h · tier 0 self-marked
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
M4 · Calibration against human judgement
0/8 · 11hreading · 1h 30m · tier 0 self-marked
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
reading · 1h · tier 0 self-marked
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
reading · 45m · tier 0 self-marked
assignment · 2h · tier 3 artifact
assignment · 3h · tier 1 machine-verified
assignment · 1h 30m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
assignment · 1h · tier 2 panel-assessed
M5 · Is the difference real?
0/6 · 8h 30mreading · 1h · tier 0 self-marked
reading · 45m · tier 0 self-marked
- optional
reading · 45m · tier 0 self-marked
Self-Consistency Improves Chain of Thought Reasoning in Language Models
assignment · 2h 30m · tier 1 machine-verified
assignment · 1h 30m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
assignment · 1h 45m · tier 1 machine-verified
M6 · Regression testing and CI
0/7 · 9h 45mlecture · 1h 30m · tier 0 self-marked
Escaping Proof-of-Concept Purgatory: Building Robust LLM Powered Applications — SciPy
reading · 1h 30m · tier 0 self-marked
reading · 1h 30m · tier 0 self-marked
reading · 1h · tier 0 self-marked
assignment · 3h · tier 3 artifact
assignment · 1h · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M7 · Course project
0/4 · 8hproject · 2h 30m · tier 2 panel-assessed
project · 2h 30m · tier 2 panel-assessed
project · 2h · tier 3 artifact
defense · 1h · tier 4 defended