ML Engineering Roadmap

Start with the pipeline and the decision behind the prediction, then follow the nine stages in order: every stage names what it needs first and what you should be able to build and defend before moving on. The stages that look boring — formulation, splitting, leakage — decide whether the interesting ones work. Progress is stored locally in your browser.

Where to start

0 / 236 lessons masteredNot started 236Learning 0Practicing 0Mastered 0
  1. 1

    The pipeline, the decision behind the prediction, and a baseline to beat

    Start here
    0/30

    What an ML system is — a pipeline from raw data to feedback — and the sixteen things that go wrong in it. Then the decision before the model: which event to predict, observed when, feeding which action, with each mistake priced. The learning paradigms and task types, and the rule, mean or majority baseline every later model must beat. Everything after this stage assumes you can name the decision and the baseline.

    Before moving on: Turn a request like "predict which customers will churn" into a formulation you can defend, say whether it needs a model at all, and produce the baseline any model you build later must beat.

  2. 2

    Datasets, splits and leakage — the part that decides whether the number means anything

    0/32

    The stage that decides whether any number you report later means anything. How a training set is built and what one row represents, which filter or join introduces which bias, the split strategies — random, temporal, grouped, stratified — and every way information from the answer leaks into the model. It ends with feature engineering and feature importance, because both are places where leakage and false causal readings enter.

    Before moving on: Choose a split from how the model will be used rather than from habit, run a leakage audit on a feature list, and read a feature-importance chart without concluding that anything causes anything.

  3. 3

    Linear models and the metrics that map to a decision

    0/17

    Linear and logistic regression as the models everything else is compared to — coefficients, residuals, the sigmoid and the threshold — then the metrics that map to a decision: the confusion matrix with a price on each cell, precision, recall, ROC AUC, PR AUC and calibration, and MAE against RMSE against MAPE for regression. It comes after splitting because a metric computed on a leaking split is a lie however well you read it.

    Before moving on: Build a confusion matrix with a price on each cell, choose a threshold from those prices rather than defaulting to 0.5, and explain to a stakeholder why high accuracy on a heavily imbalanced dataset can be worse than the majority baseline.

  4. 4

    Evaluation you can trust, bias and variance, and tree ensembles

    0/23

    Evaluation you can trust: cross-validation or a temporal holdout, sliced by the segments that matter, with an error bar on the headline number and the test set touched once. Bias and variance explain what a learning curve is telling you, and tree ensembles, k-NN, Naive Bayes and SVMs give you a set of model families to choose between with a reason. This is where "which model" becomes a question you can answer, and it needs the metrics from the stage before to answer it.

    Before moving on: Design an evaluation with slices and an error bar, read a learning curve and say whether the fix is more data, more capacity or more regularisation, and answer "which model family" with data size, latency, interpretability and cost before naming a library.

  5. 5

    Neural networks, optimisation, embeddings, architectures, foundation models and tuning

    0/34

    The mechanism every deep learning framework hides: the neuron, the forward pass, the loss, backpropagation on a computational graph, and the optimisers that walk the gradient. Then embeddings, the convolutional, recurrent and transformer architectures, pretrained foundation models and how to fine-tune them, and a tuning budget that never touches the test set. It sits after generalisation because a network with this much capacity overfits before you notice.

    Before moving on: Derive the gradients of a small network on a computational graph, explain why a learning rate that is too large diverges and one that is too small crawls, and fine-tune a pretrained model with a defensible reason for full or parameter-efficient adaptation.

  6. 6

    Time series and recommendation — where the data has an arrow

    0/12

    Two task families where the data has an arrow. Forecasting with trend, seasonality, a horizon and a rolling-origin validation that respects time, because a random split on temporal data is leakage. Recommendation as candidate generation plus ranking, collaborative and content-based signals, cold start, and the feedback loop in which the model changes the data it will be trained on next. It follows evaluation and embeddings because it applies both to data that keeps moving.

    Before moving on: Build a forecast with a validation that respects time and say which horizon it is good for, and design a recommender as candidate generation plus ranking while explaining how you would keep exploring so the feedback loop does not narrow the catalogue.

  7. 7

    Reproducible experiments, a registry, and a model that serves

    0/25

    From a notebook to a model that serves. What every run must record so it can be reproduced, versioning data, features and models, the artifact with its preprocessing inside it, the registry lifecycle, and promotion on quality, latency, memory and cost rather than one offline score. Then batch, online and streaming inference, the serving architecture, and train/serve skew — the feature engineering from stage two, reproduced identically at serving time or not.

    Before moving on: Reproduce a training run from what was recorded, choose batch, online or streaming inference from freshness, latency budget and volume, and find train/serve skew by comparing logged serving features with recomputed training features.

  8. 8

    Accelerators, distributed training, the ML test stack and responsible ML

    0/24

    What running and training a model actually costs, and what a model owes the people it affects. GPUs, memory bandwidth and VRAM, quantization, pruning and distillation, and inference cost; data, tensor and pipeline parallelism with checkpointing; the ML test stack a unit-test suite does not provide; and fairness, explainability, privacy and the line between predicting an outcome and causing it. It comes after serving because every one of these is a cost or a check on a model that already exists.

    Before moving on: Read an inference cost bill and say whether the fix is a smaller model, quantization, batching or a CPU, write the data, training, invariant and serving-contract tests a model needs, and report subgroup performance with the caveat that an explanation is approximate.

  9. 9

    Monitoring, retraining, observability, MLOps, security and system design

    0/39

    Operating a model as a system, which is where the whole domain lands. Feature, prediction and concept drift monitored separately, with drift treated as a signal to investigate rather than a trigger to retrain; champion/challenger, shadow, canary and A/B rollouts with rollback; tracing one prediction through an incident; an MLOps pipeline described without the word Kubernetes; the threats to a model; and recommendation, fraud, churn and search-ranking systems designed end to end. It needs everything before it, because each check here is on a decision made earlier.

    Before moving on: Tell data drift from concept drift from a feature bug, run a shadow deployment or a canary and roll it back, debug an incident from the business metric down to the feature that changed, and design a fraud or churn system end to end.