Skip to content

McGill ML Showcases Documentation

Welcome to the documentation site for mcgill-showcases.

This repo is organized as learning-by-doing showcase projects with reproducible scripts, clear artifacts, and track-based progression paths.

Flagship Agentic Systems Showcase

Agentic Course Assistant

Start here when you want to learn agent frameworks without making your first run depend on hosted credentials. The Agentic Course Assistant deep dive teaches routing, tools, guardrails, trace evidence, eval rubrics, and framework comparison through a deterministic offline harness first, then shows how the same design maps to optional OpenAI Agents SDK and Google ADK usage.

  • Track: Agent Frameworks
  • Project: projects/agentic-course-assistant-showcase
  • Default path: local, deterministic, and CI-safe with no API keys
  • Optional extension: live OpenAI or Google ADK runs after students install SDK extras and configure credentials

Start Here

  1. Run local setup from Getting Started (or the repo README.md).
  2. Pick a project from the catalog — or a curated track guide — based on your goal.
  3. Run one showcase end-to-end.
  4. Interpret generated artifacts with the Coverage Matrix.

Project Catalog

All 25 showcases, grouped by area and ordered so each builds on the last. Project links open the project's README.md on GitHub (the source of truth for setup and commands). For curated, guided tours of related projects see Curated Track Guides; for end-to-end sequences see Learning Paths.

Browse by area: Deep Learning Foundations · Data Foundations · Supervised Learning and Applications · Unsupervised, Causal, and NLP · Responsible AI · Optimization and Autonomous Research · Reinforcement Learning · Agents and Agentic RL · MLOps and Production Systems · Serving APIs and Observability

Deep Learning Foundations

Build the neural-network toolkit from the math up.

Data Foundations

Understand and prepare data before you model it.

Supervised Learning and Applications

Core supervised modeling and applied case studies.

Unsupervised, Causal, and NLP

Methods that go beyond labeled supervised learning.

Responsible AI

Explain, audit, and make models fair.

Optimization and Autonomous Research

Tune models and run autonomous, budgeted research loops.

  • automl-hpo-showcase — hyperparameter-optimization strategy benchmarking (grid/random/TPE).
  • autoresearch — fixed-budget autonomous research loops with Codex/Claude launch briefs.

Reinforcement Learning

Sequential decision-making from bandits to policy gradients. Track guide: Reinforcement Learning.

  • rl-bandits-policy-showcase — multi-armed bandits, reward/regret analysis, and policy recommendation.
  • student-support-rl-showcase — contextual bandits, MDPs, dynamic programming (exact Q*), Q-learning/SARSA, REINFORCE, optional DQN/PPO, reward hacking, and offline evaluation.

Agents and Agentic RL

Build agents, then learn the policies that drive them. Track guides: Agent Frameworks · Agentic RL.

  • agentic-course-assistant-showcase — agent routing, tools, guardrails, traces, eval rubrics, and optional OpenAI Agents SDK / Google ADK examples.
  • adaptive-course-assistant-rl-showcase — learned pedagogical intervention around a deterministic assistant: bandit, tutoring MDP, Q-learning/SARSA/REINFORCE, optional DQN/PPO bridge, and governance.
  • learning-agents-showcase — capstone on where learning lives in an agent: orchestration-policy RL, offline RL and off-policy evaluation, cost-aware cascades, governance, plus an OpenAI Agents SDK bridge, RLHF/DPO/GRPO/RLVR, MARL, and an optional NumPy DQN/PPO deep-RL lane.

MLOps and Production Systems

Operate, monitor, and release models safely.

Serving APIs and Observability

Ship models behind real APIs with metrics and tracing.

Curated Track Guides

Hand-picked, artifact-focused tours through related projects (each is a full page in this site):

  • Foundations — math foundations, neural-network mechanics, PyTorch training, core supervised and unsupervised workflows, EDA, and feature engineering.
  • Production — serving, drift monitoring, rollout decisions, and system patterns.
  • Ranking — grouped ranking modeling and API productization.
  • Forecasting — time-aware demand modeling and observability-ready APIs.
  • Responsible AI — fairness, explainability, and causal decision support.
  • Optimization — HPO, agentic workflows, policy optimization, reward design, and offline policy evaluation.
  • Reinforcement Learning — bandits, MDPs, dynamic programming (exact Q*), Q-learning/SARSA, REINFORCE, reward design, and an optional NumPy DQN/PPO bridge.
  • Agent Frameworks — deterministic agent workflows, course-assistant tools, guardrails, traces, eval rubrics, and OpenAI Agents SDK / Google ADK concepts.
  • Agentic RL — where learning lives inside an agent: learned intervention control, offline RL and off-policy evaluation, cost-aware cascades, RLHF/DPO/GRPO/RLVR, and multi-agent RL.

Learning Paths

For ordered, end-to-end sequences across tracks (e.g. "New to Applied ML", "Deep Learning Foundations", "ML in Production", "Learning-Agent Bridge"), see the Learning Path page. Use the Coverage Matrix to map course topics to concrete commands and artifacts.

Project Deep Dives

Contributor Entry

API Reference