Agents and Tool-Use Systems
AIE-1044 credits · 59h required · after AIE-101, AIE-102
Multi-step reliability is where demos die, and it is an arithmetic problem before it is a prompting problem. A step that succeeds 95% of the time succeeds 60% of the time across ten steps, and no amount of prompt tuning repeals that. This course starts from the arithmetic, spends most of its time on the two things that actually move it — narrower scope and better tool interfaces — and ends by attacking what you built.
The lab. Career Side Quests is the centrepiece, because it is a deterministic multi-stage pipeline that already works. You will rebuild part of it as an autonomous agent and, most likely, make it worse. That is the point: the course's second outcome is choosing between workflow and agent, and the only honest way to learn that is to pay for the wrong choice once, on a system whose correct behaviour you already know.
By the end you can
- Derive the reliability of an n-step agent from its per-step success rate, and design against it
- Choose between workflow and agent for a given task and defend the choice
- Sandbox tool access to least privilege
- Evaluate an agent on task completion rather than on transcript plausibility
- Red-team your own agent and report what you got it to do that it should not
- Justify a multi-agent design against a single-agent baseline, or decline to build one
- State what personal data enters a model, where it goes, and what the local regulator expects
0 of 49 required items complete
0m of 59h
M1 · The arithmetic of multi-step
0/5 · 5h 30mlecture · 1h · tier 0 self-marked
Building more effective AI agents — Anthropic
reading · 1h · tier 0 self-marked
lecture · 1h · tier 0 self-marked
Agentic AI MOOC | UC Berkeley CS294-196 Fall 2025 | LLM Agents Overview — Berkeley RDI
reading · 1h · tier 0 self-marked
assignment · 1h 30m · tier 1 machine-verified
M2 · Workflow or agent
0/6 · 6h 45mreading · 1h · tier 0 self-marked
reading · 45m · tier 0 self-marked
Reflexion: Language Agents with Verbal Reinforcement Learning
reading · 45m · tier 0 self-marked
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
assignment · 2h 30m · tier 1 machine-verified
assignment · 1h 30m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M3 · Tool interfaces the model can actually use
0/6 · 6h 30mreading · 1h · tier 0 self-marked
reading · 1h · tier 0 self-marked
reading · 45m · tier 0 self-marked
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
assignment · 2h 45m · tier 1 machine-verified
assignment · 45m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M4 · Reliability under repetition
0/5 · 6h 15mreading · 1h · tier 0 self-marked
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
lecture · 1h · tier 0 self-marked
LLM Agents MOOC | UC Berkeley CS294-196 Fall 2024 | LLM Reasoning — Berkeley RDI
assignment · 2h 15m · tier 1 machine-verified
assignment · 1h 45m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M5 · Sandboxing and least privilege
0/7 · 8h 45mlecture · 1h · tier 0 self-marked
LLM Agents MOOC | UC Berkeley Fall 2024 | Safe AI Agents + Evidence-based AI Policy by Dawn Song — Berkeley RDI
reading · 1h · tier 0 self-marked
reading · 1h · tier 0 self-marked
assignment · 2h 45m · tier 1 machine-verified
reading · 1h · tier 0 self-marked
Advisory Guidelines on use of Personal Data in AI Recommendation and Decision Systems
assignment · 1h 45m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M6 · Red-teaming your own agent
0/5 · 6hreading · 45m · tier 0 self-marked
Universal and Transferable Adversarial Attacks on Aligned Language Models
reading · 45m · tier 0 self-marked
WebArena: A Realistic Web Environment for Building Autonomous Agents
assignment · 2h 45m · tier 1 machine-verified
assignment · 1h 30m · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M7 · Evaluating agents
0/5 · 6hreading · 1h · tier 0 self-marked
reading · 1h · tier 0 self-marked
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
assignment · 2h 45m · tier 1 machine-verified
assignment · 1h · tier 2 panel-assessed
retention · 15m · tier 1 machine-verified
M8 · Multi-agent systems
0/6 · 6h 45mreading · 1h · tier 0 self-marked
reading · 1h 30m · tier 0 self-marked
reading · 45m · tier 0 self-marked
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
reading · 45m · tier 0 self-marked
assignment · 2h 30m · tier 1 machine-verified
retention · 15m · tier 1 machine-verified
M9 · Course project
0/4 · 6h 30mproject · 2h 45m · tier 3 artifact
project · 1h 30m · tier 2 panel-assessed
project · 1h 15m · tier 3 artifact
defense · 1h · tier 4 defended