← Role tracks

R3 Stage D · 160 h (48 T / 112 P)

MLOps / LLMOps Engineer

Owns the path to production and the reliability of AI systems in operation.

Tech stack

DockerKubernetesHelmMLflowDVCAirflowTerraformGitHub ActionsPrometheusGrafanaEvidently

Modules

M1 · Containerisation and orchestration

40 h
TOPICS
  • Image design and minimisation
  • Multi-stage builds
  • Registries and scanning
  • Kubernetes objects: pod, deployment, service, ingress, config, secret
  • Resource requests and limits
  • Autoscaling
  • GPU scheduling
  • Node pools
  • Rollout strategies

Lab: Deploy a scalable, resource-bounded model service on Kubernetes with autoscaling

Course material for this module is in production.

M2 · Reproducibility and lineage

40 h
TOPICS
  • Experiment tracking
  • Model registry and staging transitions
  • Dataset and artefact versioning
  • Deterministic environments
  • Orchestration DAGs
  • Scheduling and backfills
  • Dependency and failure handling

Lab: Automated, versioned training-to-registry pipeline with scheduled retraining

Course material for this module is in production.

M3 · Continuous delivery for AI

40 h
TOPICS
  • Test strategy for ML and LLM systems
  • Evaluation gates in CI/CD
  • Infrastructure as code
  • Environment promotion
  • Canary and blue-green deployment
  • Automated rollback
  • Secrets and supply-chain security

Lab: Pipeline that automatically blocks promotion when evaluation scores regress

Course material for this module is in production.

M4 · Operations and observability

40 h
TOPICS
  • Metrics, logs and traces
  • Data, concept and prediction drift
  • Performance and cost dashboards
  • Alerting and thresholds
  • SLO and error-budget management
  • Incident response and post-mortems
  • Capacity planning

Lab: Drift detection with alerting plus an incident runbook and simulated incident

Course material for this module is in production.

Track project

Production ML platform: pipeline, registry, evaluation-gated CI/CD, canary deployment with automated rollback, drift monitoring and alerting, cost dashboard, and an operational runbook — demonstrated through an induced failure and recovery.

Job-ready exit standard

Can deploy, monitor and version models in production; cloud associate or ML certification recommended.