Indie Degree
← Programme

Inference, Cost and Latency Engineering

AIE-105

3 credits · 50h required · 2h optional · after AIE-101

The two questions a field engineer is actually asked in the room are what this will cost at our volume and how fast it will feel. Very few candidates can answer either with a number, and the ones who can are visibly different in the meeting. This course exists to make you one of them.

The lab. Your own products, and one of them is unusually well suited: 1 Percent More Fluent streams both text and audio, so it has a real time-to-first-token problem with a real user on the other end. Career Side Quests is the cost case — a multi-stage pipeline with per-run accounting already in place.

By the end you can

  • Model the token cost of a feature at a given volume, before building it
  • Explain the prefill/decode split and what each implies for latency
  • Apply caching, batching and streaming, and measure what each bought
  • Read a serving stack — paged attention, continuous batching, quantisation — and reason about its trade-offs
  • Degrade a model-backed feature gracefully rather than failing it

0 of 42 required items complete

0m of 50h 15m

M1 · The arithmetic of cost

0/6 · 8h 45m
  • reading · 1h 30m · tier 0 self-marked

    Transformer Inference Arithmetic

  • lecture · 1h 30m · tier 0 self-marked

    Stanford CS336 Language Modeling from Scratch I 2025

  • reading · 45m · tier 0 self-marked

    AI Engineering

  • exemptable

    assignment · 1h · tier 1 machine-verified

  • assignment · 2h · tier 1 machine-verified

  • optional

    assignment · 1h 45m · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M2 · Prefill and decode

0/5 · 6h 45m
  • reading · 1h 30m · tier 0 self-marked

    Efficiently Scaling Transformer Inference

  • reading · 1h · tier 0 self-marked

    Making Deep Learning Go Brrrr From First Principles

  • assignment · 2h 30m · tier 1 machine-verified

  • assignment · 1h 30m · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M3 · Caching

0/5 · 5h 30m
  • reading · 45m · tier 0 self-marked

    Prompt caching

  • reading · 1h · tier 0 self-marked

    BerriAI/litellm

  • assignment · 2h 30m · tier 1 machine-verified

  • assignment · 1h · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M4 · Batching and serving

0/6 · 7h 30m

M5 · Quantisation

0/7 · 6h 30m

M6 · Streaming and perceived latency

0/5 · 5h 30m

M7 · Reliability

0/4 · 4h 45m
  • reading · 1h 30m · tier 0 self-marked

    Release It!

  • reading · 45m · tier 0 self-marked

    BerriAI/litellm

  • assignment · 2h 15m · tier 1 machine-verified

  • retention · 15m · tier 1 machine-verified

M8 · Course project

0/4 · 6h 45m
  • project · 2h 45m · tier 3 artifact

  • project · 1h 45m · tier 2 panel-assessed

  • project · 1h 15m · tier 3 artifact

  • defense · 1h · tier 4 defended