ML Engineering Roadmap
Start with the pipeline and the decision behind the prediction, then follow the nine stages in order: every stage names what it needs first and what you should be able to build and defend before moving on. The stages that look boring — formulation, splitting, leakage — decide whether the interesting ones work. Progress is stored locally in your browser.
Where to start
Machine Learning Engineering
9 stages · 0/236 lessonsTurning data into a model that keeps working after it ships: formulation, data, evaluation, training, serving and operation.
- The pipeline, the decision behind the prediction, and a baseline to beat
- Datasets, splits and leakage — the part that decides whether the number means anything
- Linear models and the metrics that map to a decision
- Evaluation you can trust, bias and variance, and tree ensembles
- Neural networks, optimisation, embeddings, architectures, foundation models and tuning
- Time series and recommendation — where the data has an arrow
- Reproducible experiments, a registry, and a model that serves
- Accelerators, distributed training, the ML test stack and responsible ML
- Monitoring, retraining, observability, MLOps, security and system design
- 10/30
The pipeline, the decision behind the prediction, and a baseline to beat
Start hereWhat an ML system is — a pipeline from raw data to feedback — and the sixteen things that go wrong in it. Then the decision before the model: which event to predict, observed when, feeding which action, with each mistake priced. The learning paradigms and task types, and the rule, mean or majority baseline every later model must beat. Everything after this stage assumes you can name the decision and the baseline.
Before moving on: Turn a request like "predict which customers will churn" into a formulation you can defend, say whether it needs a model at all, and produce the baseline any model you build later must beat.
- What ML Engineering Is
- The ML Pipeline
- What Can Go Wrong
- Learning vs Programming
- The ML Reasoning Loop
- Don't Delegate Understanding
- Problem Formulation
- Decision Before Model
- Prediction vs Decision
- Target Definition
- Label Construction
- Label Leakage
- When Not to Use ML
- Supervised Learning
- Unsupervised Learning
- Semi-Supervised Learning
- Self-Supervised Learning
- The Learning Signal
- Choosing a Paradigm
- Regression
- Classification
- Ranking
- Clustering
- Dimensionality Reduction
- Anomaly Detection
- Baselines Are Mandatory
- The Rule Baseline
- Majority Class and Mean Predictor
- The Linear Baseline
- Beating the Baseline
- 20/32
Datasets, splits and leakage — the part that decides whether the number means anything
The stage that decides whether any number you report later means anything. How a training set is built and what one row represents, which filter or join introduces which bias, the split strategies — random, temporal, grouped, stratified — and every way information from the answer leaks into the model. It ends with feature engineering and feature importance, because both are places where leakage and false causal readings enter.
Before moving on: Choose a split from how the model will be used rather than from habit, run a leakage audit on a feature list, and read a feature-importance chart without concluding that anything causes anything.
- Dataset Construction
- What Is One Example?
- Sampling Strategies
- Selection Bias
- Survivorship Bias
- Class Imbalance
- Label Quality
- Train / Validation / Test
- Random Split
- Time-Based Split
- Group Split
- Stratified Split
- Choosing a Split Strategy
- Data Leakage
- Target Leakage
- Temporal Leakage
- Entity Leakage
- Preprocessing Leakage
- Evaluation Leakage
- The Leakage Audit
- Feature Engineering
- Aggregation Features
- Bucketing & Normalisation
- Categorical Encoding
- Target Encoding
- Temporal Features
- Missing Data
- Raw Features vs Learned Representations
- Feature Selection
- Feature Importance
- Permutation Importance
- Attribution Is Not Causality
- 30/17
Linear models and the metrics that map to a decision
Linear and logistic regression as the models everything else is compared to — coefficients, residuals, the sigmoid and the threshold — then the metrics that map to a decision: the confusion matrix with a price on each cell, precision, recall, ROC AUC, PR AUC and calibration, and MAE against RMSE against MAPE for regression. It comes after splitting because a metric computed on a leaking split is a lie however well you read it.
Before moving on: Build a confusion matrix with a price on each cell, choose a threshold from those prices rather than defaulting to 0.5, and explain to a stakeholder why high accuracy on a heavily imbalanced dataset can be worse than the majority baseline.
- Linear Regression
- Residuals & Assumptions
- Logistic Regression
- Sigmoid & Probability
- Thresholding
- Regularized Linear Models
- The Confusion Matrix
- Precision, Recall & F1
- Threshold Selection
- Accuracy Under Imbalance
- ROC AUC
- PR AUC
- Calibration
- MSE, RMSE and MAE
- R² (Coefficient of Determination)
- MAPE and Its Caveats
- Choosing a Regression Metric
- 40/23
Evaluation you can trust, bias and variance, and tree ensembles
Evaluation you can trust: cross-validation or a temporal holdout, sliced by the segments that matter, with an error bar on the headline number and the test set touched once. Bias and variance explain what a learning curve is telling you, and tree ensembles, k-NN, Naive Bayes and SVMs give you a set of model families to choose between with a reason. This is where "which model" becomes a question you can answer, and it needs the metrics from the stage before to answer it.
Before moving on: Design an evaluation with slices and an error bar, read a learning curve and say whether the fix is more data, more capacity or more regularisation, and answer "which model family" with data size, latency, interpretability and cost before naming a library.
Needs first:Datasets, splits and leakage — the part that decides whether the number means anythingLinear models and the metrics that map to a decision- Business Metrics vs Model Metrics
- Offline vs Online Evaluation
- Cross-Validation
- Time-Series Validation
- Evaluation Slices
- Metric Uncertainty
- Never Tune on the Test Set
- Bias and Variance
- Overfitting
- Underfitting
- Learning Curves
- Regularisation
- Early Stopping
- Decision Trees
- How a Tree Chooses a Split
- Random Forests
- Gradient Boosting
- XGBoost and LightGBM as Implementations
- Tree Ensembles: When and When Not
- k-Nearest Neighbours
- Naive Bayes
- Support Vector Machines
- Which Model Should We Use?
- 50/34
Neural networks, optimisation, embeddings, architectures, foundation models and tuning
The mechanism every deep learning framework hides: the neuron, the forward pass, the loss, backpropagation on a computational graph, and the optimisers that walk the gradient. Then embeddings, the convolutional, recurrent and transformer architectures, pretrained foundation models and how to fine-tune them, and a tuning budget that never touches the test set. It sits after generalisation because a network with this much capacity overfits before you notice.
Before moving on: Derive the gradients of a small network on a computational graph, explain why a learning rate that is too large diverges and one that is too small crawls, and fine-tune a pretrained model with a defensible reason for full or parameter-efficient adaptation.
- Neural Networks
- The Neuron
- Activation Functions
- The Forward Pass
- Loss Functions
- Backpropagation
- Computational Graphs
- Gradient Descent
- Optimisers: SGD, Momentum, Adam
- Epoch, Batch, Step
- Batch Size and Learning Rate
- Vanishing and Exploding Gradients
- Normalisation Layers
- Initialisation and Convergence
- Embeddings
- Embedding Training
- Cosine Similarity
- Embedding Projection Caveats
- Embedding Drift
- CNN Concepts
- Sequence Models
- Transformer Fundamentals
- Self-Attention
- Positional Information
- Encoder / Decoder Families
- Foundation Models
- Transfer Learning
- Fine-Tuning
- Parameter-Efficient Fine-Tuning
- Hyperparameters
- Grid Search and Random Search
- Bayesian Optimisation, Successive Halving and Early Termination
- The Tuning Budget
- AutoML Hides the Search
- 60/12
Time series and recommendation — where the data has an arrow
Two task families where the data has an arrow. Forecasting with trend, seasonality, a horizon and a rolling-origin validation that respects time, because a random split on temporal data is leakage. Recommendation as candidate generation plus ranking, collaborative and content-based signals, cold start, and the feedback loop in which the model changes the data it will be trained on next. It follows evaluation and embeddings because it applies both to data that keeps moving.
Before moving on: Build a forecast with a validation that respects time and say which horizon it is good for, and design a recommender as candidate generation plus ranking while explaining how you would keep exploring so the feedback loop does not narrow the catalogue.
- 70/25
Reproducible experiments, a registry, and a model that serves
From a notebook to a model that serves. What every run must record so it can be reproduced, versioning data, features and models, the artifact with its preprocessing inside it, the registry lifecycle, and promotion on quality, latency, memory and cost rather than one offline score. Then batch, online and streaming inference, the serving architecture, and train/serve skew — the feature engineering from stage two, reproduced identically at serving time or not.
Before moving on: Reproduce a training run from what was recorded, choose batch, online or streaming inference from freshness, latency budget and volume, and find train/serve skew by comparing logged serving features with recomputed training features.
Needs first:Datasets, splits and leakage — the part that decides whether the number means anythingEvaluation you can trust, bias and variance, and tree ensembles- Experiment Tracking
- Reproducibility
- Random Seeds
- Dataset Versioning
- Feature and Model Versioning
- Model Lineage
- What a Model Artifact Contains
- The Model Registry
- Promotion Is a Checklist, Not a Score
- Artifact Integrity
- Preprocessing Lives in the Artifact
- Batch Inference
- Online Inference
- Streaming Inference
- Choosing the Inference Mode
- Model Serving Architecture
- Inference Batching
- CPU or GPU for Inference
- Train / Serve Skew
- Feature Stores
- Point-in-Time Correctness
- Feature Freshness
- Latency Breakdown
- Throughput vs Latency
- Serving Fallbacks
- 80/24
Accelerators, distributed training, the ML test stack and responsible ML
What running and training a model actually costs, and what a model owes the people it affects. GPUs, memory bandwidth and VRAM, quantization, pruning and distillation, and inference cost; data, tensor and pipeline parallelism with checkpointing; the ML test stack a unit-test suite does not provide; and fairness, explainability, privacy and the line between predicting an outcome and causing it. It comes after serving because every one of these is a cost or a check on a model that already exists.
Before moving on: Read an inference cost bill and say whether the fix is a smaller model, quantization, batching or a CPU, write the data, training, invariant and serving-contract tests a model needs, and report subgroup performance with the caveat that an explanation is approximate.
Needs first:Neural networks, optimisation, embeddings, architectures, foundation models and tuningReproducible experiments, a registry, and a model that serves- GPU Fundamentals
- Memory Bandwidth & VRAM
- Quantization
- Model Compression
- Pruning & Distillation
- Inference Cost
- Distributed Training
- Data Parallelism
- Model, Tensor & Pipeline Parallelism
- Gradient Synchronisation
- Checkpointing
- Training Cost
- The ML Testing Stack
- Data & Feature Tests
- Training Smoke Tests
- Model Invariant Tests
- Serving Contract Tests
- Robustness Testing
- Model Regression Tests
- Fairness
- Explainability
- Causality vs Prediction
- ML Privacy
- Human Oversight
- 90/39
Monitoring, retraining, observability, MLOps, security and system design
Operating a model as a system, which is where the whole domain lands. Feature, prediction and concept drift monitored separately, with drift treated as a signal to investigate rather than a trigger to retrain; champion/challenger, shadow, canary and A/B rollouts with rollback; tracing one prediction through an incident; an MLOps pipeline described without the word Kubernetes; the threats to a model; and recommendation, fraud, churn and search-ranking systems designed end to end. It needs everything before it, because each check here is on a decision made earlier.
Before moving on: Tell data drift from concept drift from a feature bug, run a shadow deployment or a canary and roll it back, debug an incident from the business metric down to the feature that changed, and design a fraud or churn system end to end.
Needs first:Reproducible experiments, a registry, and a model that servesAccelerators, distributed training, the ML test stack and responsible ML- Model Monitoring
- Data Drift
- Feature Drift
- Prediction Drift
- Concept Drift
- Drift Is Not Failure
- Ground-Truth Delay
- Performance Decay
- Retraining as a Decision
- Retraining Strategies
- Champion / Challenger
- Shadow Deployment
- Canary Rollout
- A/B Testing Models
- Rollback & Fallback
- ML Observability
- Prediction Logging
- Tracing a Prediction
- ML Incident Debugging
- Model Postmortems
- What MLOps Is
- The MLOps Pipeline
- CI, CT and CD for ML
- ML Platform Engineering
- Cloud ML Services
- ML Orchestration
- ML Cost Optimisation
- ML Security
- Data Poisoning
- Adversarial Inputs
- The Model Supply Chain
- Inference Abuse
- ML System Design Architecture
- The Questions Before the Boxes
- Designing a Recommendation System
- Designing a Fraud Detection System
- Designing a Churn Prediction System
- Designing Search Ranking
- The Boundary With Agentic Engineering