Indie Degree
← Programme

Retrieval and Context Systems

AIE-103

3 credits · 50h required · 1h optional · after AIE-101

Retrieval is the default answer to most enterprise AI questions and is wrong often enough that knowing when not to use it is the differentiator. This course refuses the word RAG for as long as possible, because it welds together two separate problems — finding the right thing, and writing a good answer from it — and most broken systems are broken in only one of them.

The lab. Two corpora of your own, chosen because they fail differently. Carpark SG's official rate documents are dense with codes, acronyms and proper nouns — HDB and URA terminology that embeddings blur together and lexical search nails. Indie Degree's own resource corpus is prose-heavy with near-synonymous topics, where lexical search fails and embeddings shine. If a technique only helps on one of them, that is the finding.

By the end you can

  • Build hybrid retrieval and show where lexical beats dense, with numbers
  • Choose a chunking strategy from measured retrieval quality rather than from a blog post
  • Evaluate retrieval separately from generation, and locate which half is failing
  • Articulate when long context replaces retrieval and when it does not
  • Enforce document-level access control through a retrieval pipeline, and prove it holds

0 of 40 required items complete

0m of 50h 15m

M1 · Two problems wearing one name

0/5 · 5h 45m

M2 · Embeddings and the limits of similarity

0/5 · 6h

M3 · Chunking, measured

0/5 · 6h 15m
  • reading · 1h · tier 0 self-marked

    Introducing Contextual Retrieval

  • reading · 45m · tier 0 self-marked

    Searching for Best Practices in Retrieval-Augmented Generation

  • assignment · 2h 30m · tier 1 machine-verified

  • assignment · 1h 45m · tier 1 machine-verified

  • retention · 15m · tier 1 machine-verified

M4 · Hybrid and lexical search

0/5 · 6h 15m

M5 · Reranking

0/5 · 5h
  • reading · 45m · tier 0 self-marked

    Pretrained Models — Cross Encoder

  • reading · 45m · tier 0 self-marked

    Precise Zero-Shot Dense Retrieval without Relevance Labels

  • assignment · 2h · tier 1 machine-verified

  • assignment · 1h 15m · tier 2 panel-assessed

  • retention · 15m · tier 1 machine-verified

M6 · Evaluating retrieval separately from generation

0/5 · 5h 30m
  • reading · 45m · tier 0 self-marked

    RAGAS: Automated Evaluation of Retrieval Augmented Generation

  • reading · 1h · tier 0 self-marked

    explodinggradients/ragas

  • assignment · 2h · tier 1 machine-verified

  • assignment · 1h 30m · tier 1 machine-verified

  • retention · 15m · tier 1 machine-verified

M7 · When long context replaces retrieval

0/2 · 4h 15m

M8 · Production retrieval

0/4 · 5h 30m
  • reading · 45m · tier 0 self-marked

    OWASP Top 10 for Large Language Model Applications

  • assignment · 2h 30m · tier 1 machine-verified

  • assignment · 2h · tier 1 machine-verified

  • retention · 15m · tier 1 machine-verified

M9 · Course project

0/4 · 6h 30m
  • project · 2h 30m · tier 3 artifact

  • project · 1h 45m · tier 2 panel-assessed

  • project · 1h 15m · tier 3 artifact

  • defense · 1h · tier 4 defended