R3 Stage D · 160 h (48 T / 112 P)
MLOps / LLMOps Engineer
Owns the path to production and the reliability of AI systems in operation.
Tech stack
DockerKubernetesHelmMLflowDVCAirflowTerraformGitHub ActionsPrometheusGrafanaEvidently
Modules
M1 · Containerisation and orchestration
40 h TOPICS
- Image design and minimisation
- Multi-stage builds
- Registries and scanning
- Kubernetes objects: pod, deployment, service, ingress, config, secret
- Resource requests and limits
- Autoscaling
- GPU scheduling
- Node pools
- Rollout strategies
Lab: Deploy a scalable, resource-bounded model service on Kubernetes with autoscaling
Course material for this module is in production.
M2 · Reproducibility and lineage
40 h TOPICS
- Experiment tracking
- Model registry and staging transitions
- Dataset and artefact versioning
- Deterministic environments
- Orchestration DAGs
- Scheduling and backfills
- Dependency and failure handling
Lab: Automated, versioned training-to-registry pipeline with scheduled retraining
Course material for this module is in production.
M3 · Continuous delivery for AI
40 h TOPICS
- Test strategy for ML and LLM systems
- Evaluation gates in CI/CD
- Infrastructure as code
- Environment promotion
- Canary and blue-green deployment
- Automated rollback
- Secrets and supply-chain security
Lab: Pipeline that automatically blocks promotion when evaluation scores regress
Course material for this module is in production.
M4 · Operations and observability
40 h TOPICS
- Metrics, logs and traces
- Data, concept and prediction drift
- Performance and cost dashboards
- Alerting and thresholds
- SLO and error-budget management
- Incident response and post-mortems
- Capacity planning
Lab: Drift detection with alerting plus an incident runbook and simulated incident
Course material for this module is in production.
Track project
Production ML platform: pipeline, registry, evaluation-gated CI/CD, canary deployment with automated rollback, drift monitoring and alerting, cost dashboard, and an operational runbook — demonstrated through an induced failure and recovery.
Job-ready exit standard
Can deploy, monitor and version models in production; cloud associate or ML certification recommended.