Inference, Cost and Latency Engineering
AIE-1053 credits · 50h required · 2h optional · after AIE-101
The two questions a field engineer is actually asked in the room are what this will cost at our volume and how fast it will feel. Very few candidates can answer either with a number, and the ones who can are visibly different in the meeting. This course exists to make you one of them.
The lab. Your own products, and one of them is unusually well suited: 1 Percent More Fluent streams both text and audio, so it has a real time-to-first-token problem with a real user on the other end. Career Side Quests is the cost case — a multi-stage pipeline with per-run accounting already in place.
By the end you can
- Model the token cost of a feature at a given volume, before building it
- Explain the prefill/decode split and what each implies for latency
- Apply caching, batching and streaming, and measure what each bought
- Read a serving stack — paged attention, continuous batching, quantisation — and reason about its trade-offs
- Degrade a model-backed feature gracefully rather than failing it
0 of 42 required items complete
0m of 50h 15m
M1 · The arithmetic of cost
0/6 · 8h 45mreading · 1h 30m · tier 0 self-marked
lecture · 1h 30m · tier 0 self-marked
reading · 45m · tier 0 self-marked
- exemptable
assignment · 1h · tier 1 machine-verified
assignment · 2h · tier 1 machine-verified
- optional
assignment · 1h 45m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M2 · Prefill and decode
0/5 · 6h 45mreading · 1h 30m · tier 0 self-marked
reading · 1h · tier 0 self-marked
assignment · 2h 30m · tier 1 machine-verified
assignment · 1h 30m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M3 · Caching
0/5 · 5h 30mreading · 45m · tier 0 self-marked
reading · 1h · tier 0 self-marked
assignment · 2h 30m · tier 1 machine-verified
assignment · 1h · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M4 · Batching and serving
0/6 · 7h 30mreading · 1h · tier 0 self-marked
Efficient Memory Management for Large Language Model Serving with PagedAttention
reading · 1h · tier 0 self-marked
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
reading · 1h · tier 0 self-marked
assignment · 3h · tier 1 machine-verified
assignment · 1h 15m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M5 · Quantisation
0/7 · 6h 30mreading · 45m · tier 0 self-marked
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
reading · 45m · tier 0 self-marked
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
reading · 45m · tier 0 self-marked
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
reading · 45m · tier 0 self-marked
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
reading · 45m · tier 0 self-marked
assignment · 2h 30m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M6 · Streaming and perceived latency
0/5 · 5h 30mreading · 45m · tier 0 self-marked
reading · 45m · tier 0 self-marked
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
assignment · 2h 30m · tier 1 machine-verified
assignment · 1h 15m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M7 · Reliability
0/4 · 4h 45mreading · 1h 30m · tier 0 self-marked
reading · 45m · tier 0 self-marked
assignment · 2h 15m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M8 · Course project
0/4 · 6h 45mproject · 2h 45m · tier 3 artifact
project · 1h 45m · tier 2 panel-assessed
project · 1h 15m · tier 3 artifact
defense · 1h · tier 4 defended