Indie Degree
← Programme

Speech and Multimodal Systems

AIE-106

4 credits · 62h required · after AIE-101, AIE-105

Voice is the domain where the target employers live, and it is unusually unforgiving: a text product that takes two seconds is fine, and a voice product that takes two seconds is broken. Human conversation runs on a 300–500 ms response window, and most deployed voice agents sit at 1.4 seconds or worse. This course is about closing that gap and knowing exactly which hop to blame.

The lab. 1 Percent More Fluent and LearnIndo. Both already move audio in one direction; neither closes the loop, measures what it costs, or handles a user interrupting.

By the end you can

  • Build a full ASR → reasoning → TTS loop and account for every millisecond in the budget
  • Explain why transport choice dominates perceived latency in voice agents
  • Measure ASR quality properly, including on accented and code-switched speech
  • Handle barge-in, turn-taking and partial results
  • Decide whether a document problem needs vision at all, and show the measurement behind the answer

0 of 53 required items complete

0m of 62h 30m

M1 · How speech becomes tokens

0/8 · 7h 30m

M2 · Measuring recognition properly

0/6 · 7h

M3 · Accented and code-switched speech

0/6 · 7h 30m

M4 · Synthesis

0/8 · 8h 15m

M5 · Transport and streaming

0/5 · 6h 30m
  • reading · 1h 30m · tier 0 self-marked

    WebRTC for the Curious

  • reading · 1h · tier 0 self-marked

    Understand and improve voice agent latency

  • assignment · 2h 45m · tier 1 machine-verified

  • assignment · 1h · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M6 · Turn-taking and barge-in

0/5 · 6h 30m

M7 · Documents and vision-language models

0/11 · 12h 30m

M8 · Course project

0/4 · 6h 45m
  • project · 2h 45m · tier 3 artifact

  • project · 1h 45m · tier 2 panel-assessed

  • project · 1h 15m · tier 3 artifact

  • defense · 1h · tier 4 defended