Indie Degree
← Programme

Evaluation and Measurement

AIE-102

4 credits · 66h required · 1h optional · after AIE-101

You have shipped four LLM products and cannot currently answer, with a number, whether any of them is any good. This course fixes that, using those products as the laboratory. Nothing here is a toy dataset.

By the end you can

  • Build an eval set from production failures rather than from imagination
  • Name and measure the known pathologies of LLM judges — position, verbosity, self-enhancement bias
  • Calibrate an automated judge against your own labels and report the agreement, not just the score
  • Decide whether a difference between two systems is real, with an interval rather than a vibe
  • Wire evals into CI so a regression fails a build

0 of 47 required items complete

0m of 66h 10m

M1 · What an eval is, and what it is not

0/6 · 6h 40m

M2 · Error analysis and the first eval set

0/6 · 9h
  • lecture · 1h 30m · tier 0 self-marked

    How to Build and Evaluate AI systems in the Age of LLMsDataTalksClub ⬛

  • assignment · 3h · tier 2 panel-assessed

  • assignment · 3h · tier 1 machine-verified

  • retention · 15m · tier 1 machine-verified

  • assignment · 30m · tier 2 panel-assessed

  • reading · 45m · tier 0 self-marked

    LLM Evals: Everything You Need to Know

M3 · LLM-as-judge and its pathologies

0/10 · 14h

M4 · Calibration against human judgement

0/8 · 11h

M5 · Is the difference real?

0/6 · 8h 30m

M6 · Regression testing and CI

0/7 · 9h 45m

M7 · Course project

0/4 · 8h
  • project · 2h 30m · tier 2 panel-assessed

  • project · 2h 30m · tier 2 panel-assessed

  • project · 2h · tier 3 artifact

  • defense · 1h · tier 4 defended