Indie Degree
← Programme

Agents and Tool-Use Systems

AIE-104

4 credits · 59h required · after AIE-101, AIE-102

Multi-step reliability is where demos die, and it is an arithmetic problem before it is a prompting problem. A step that succeeds 95% of the time succeeds 60% of the time across ten steps, and no amount of prompt tuning repeals that. This course starts from the arithmetic, spends most of its time on the two things that actually move it — narrower scope and better tool interfaces — and ends by attacking what you built.

The lab. Career Side Quests is the centrepiece, because it is a deterministic multi-stage pipeline that already works. You will rebuild part of it as an autonomous agent and, most likely, make it worse. That is the point: the course's second outcome is choosing between workflow and agent, and the only honest way to learn that is to pay for the wrong choice once, on a system whose correct behaviour you already know.

By the end you can

  • Derive the reliability of an n-step agent from its per-step success rate, and design against it
  • Choose between workflow and agent for a given task and defend the choice
  • Sandbox tool access to least privilege
  • Evaluate an agent on task completion rather than on transcript plausibility
  • Red-team your own agent and report what you got it to do that it should not
  • Justify a multi-agent design against a single-agent baseline, or decline to build one
  • State what personal data enters a model, where it goes, and what the local regulator expects

0 of 49 required items complete

0m of 59h

M1 · The arithmetic of multi-step

0/5 · 5h 30m

M2 · Workflow or agent

0/6 · 6h 45m

M3 · Tool interfaces the model can actually use

0/6 · 6h 30m

M4 · Reliability under repetition

0/5 · 6h 15m

M5 · Sandboxing and least privilege

0/7 · 8h 45m

M6 · Red-teaming your own agent

0/5 · 6h

M7 · Evaluating agents

0/5 · 6h
  • reading · 1h · tier 0 self-marked

    UKGovernmentBEIS/inspect_ai

  • reading · 1h · tier 0 self-marked

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

  • assignment · 2h 45m · tier 1 machine-verified

  • assignment · 1h · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M8 · Multi-agent systems

0/6 · 6h 45m

M9 · Course project

0/4 · 6h 30m
  • project · 2h 45m · tier 3 artifact

  • project · 1h 30m · tier 2 panel-assessed

  • project · 1h 15m · tier 3 artifact

  • defense · 1h · tier 4 defended