Module 2: Pipelines, orchestration, and quality#

Theme#

Pipelines, orchestration, and quality

Essential Question#

How do data pipelines fail, and how are failures detected?

Module Components#

  • Book prose: conceptual framing, domain scenario, methods, and failure modes

  • Assignment: evidence-backed production of a specific artifact

  • Slides: presentation sequence for seminar or lecture delivery

  • Narration: spoken version of the slide flow

  • Rubric: criteria for evaluating the module artifact

  • Notebook: executable lab aligned with the module theme using synthetic pipeline events with freshness, schema drift, lineage completeness, volume, and access-risk indicators

Module Artifact#

AI data platform design review with lineage, quality checks, cost controls, and access model focused on pipelines, orchestration, and quality: Build a small quality-checked pipeline.

Professional Setting#

Students work as if advising a platform team designing data infrastructure for repeatable model training and monitoring. Their work must be intelligible to data engineer, ML engineer, security architect, data steward, and platform owner.

Use This Module in Order#

  1. Read the learning chapter.

  2. Review the slide deck with the matching narration.

  3. In Populi, open the private student-repository link for this course and enter modules/module-2.

  4. Clone the repository once or open its Codespace/Colab copy; run lab.ipynb and complete exercise.ipynb there.

  5. Self-check with the rubric, commit and push the work, then submit exactly what Populi requests.