DevOps Roadmap

Nine stages in one order, starting at Source to artifact. Every stage names what it needs first and what you should be able to safely do before moving on — the order is load-bearing, because a practice built on an unstable artifact cannot be made safe further down the pipeline. Progress is stored locally in your browser.

Where to start

0 / 185 lessons masteredNot started 185Learning 0Practicing 0Mastered 0
  1. 1

    Source to artifact

    Start here
    0/18

    Where the loop begins: what makes production different, why a commit is a candidate for production state rather than history, and how CI, builds and registries turn that commit into one immutable, addressable artifact. It comes first because everything after it assumes the artifact is the unit that moves.

    Before moving on: Answer "what exactly would I deploy, and which commit produced it" with a digest, and explain why rebuilding per environment destroys that evidence.

  2. 2

    A runnable unit and the things that vary around it

    0/19

    Packaging the artifact so it runs the same way everywhere, then supplying the parts that legitimately differ per environment — configuration and credentials — without rebuilding it. Container layers and the process and signal model, config validated at startup, workload identity instead of pasted secrets, and why staging is not production.

    Before moving on: Ship artifact plus config instead of "the code": build a small image, say what happens to it on SIGTERM, and place each value in the image, the config or the secret store.

    Needs first:Source to artifact
  3. 3

    Getting change into production safely

    0/18

    Deployment stops being an event and becomes a routine. Release versus deployment, the pipeline that carries one artifact through environments, the strategies — recreate, rolling, blue/green, canary, flags — with what each costs and how you get back, and blast radius as the idea that organises all of it. It needs a stable artifact and config to move, which is why it sits after the first two stages.

    Before moving on: Put a new version in front of real traffic in a shape you chose, explain why version coexistence is the hard part, and decide whether rolling back or rolling forward is safe for a given change.

  4. 4

    Infrastructure you can reproduce

    0/13

    The environment the artifact runs in becomes code that can be reviewed, planned and diffed rather than a console someone clicked. Plans, state, drift, the destructive change a rename hides, immutable infrastructure and preview environments — and, because the console is still there, who is allowed to touch production by hand and with what privilege.

    Before moving on: Stand an environment up from a repository, read a plan and say what it will destroy before it does, and detect when reality has stopped matching the code.

  5. 5

    Orchestration, discovery and traffic

    0/20

    Running many replicas of many services, letting them find each other, and putting traffic in front of them without dropping requests. Kubernetes is taught as one implementation of the scheduling and reconciliation problem — pods, deployments, requests and limits, probes — and then the operational half of the network: DNS under change, load balancer health, connection draining and certificate lifecycles.

    Before moving on: Explain reconciliation and the scheduler well enough to diagnose an OOM kill, a failing probe or a pending pod, and drain a replica out of a load balancer without losing a request.

  6. 6

    Operating it, and responding when it breaks

    0/20

    What an operator does with observability: alerts that page only for symptoms a human must act on, dashboards annotated with deploys, on-call that stays healthy, and the incident loop from detection through mitigation to a postmortem whose action items change the system. It sits here because you need something deployed and running before there is anything to operate — and because "what changed?" is only answerable once deploys leave a trail.

    Before moving on: Tell whether production is healthy without asking anybody, run an incident from detection to mitigation, and write a postmortem whose action items change the system rather than the people.

  7. 7

    Capacity, cost and the stateful parts

    0/27

    How much traffic the system can take, what happens when it takes more, what each unit of work costs, and how to change a schema without an outage. Headroom and load shedding, autoscaling signals and their failure modes, cost drivers, and expand/migrate/contract with the backfills and locks that make a migration the change most likely to cause an outage. Reliability, money and the database stop being separate conversations.

    Before moving on: Build a capacity model with headroom for failure and deploys, pick an autoscaling signal that reflects the real constraint, and run a zero-downtime migration as one event coupled to its deploy.

  8. 8

    Release engineering, supply chain and platform

    0/25

    Delivery becomes a product with users. Release manifests, change management and an audit trail that answers "what is in production, where did it come from and who approved it"; provenance, signing, SBOMs and scanning for everything between a dependency and a running artifact; and golden paths, guardrails and policy as code that make the safe path the easy one. It builds on the artifact, the pipeline and the infrastructure code of the earlier stages.

    Before moving on: Answer what is in production with evidence rather than recollection, verify an artifact's provenance before promoting it, and design a golden path other teams choose over doing it by hand.

  9. 9

    Multi-region, recovery and production engineering

    0/25

    Losing a region, a database or a dependency and still having a tested procedure back to service. Backups you have actually restored, RTO and RPO tied to real runbooks, region failover and the capacity question it raises, readiness reviews and service ownership, break-glass access, and the same discipline applied to agent systems, where prompts and models are deployable inputs. Last because it engineers the properties that let an organisation keep many services up, not one.

    Before moving on: Run a restore drill and a region failover inside objectives you have measured, review a service for readiness before it carries traffic, and roll out a prompt or model change behind a canary with a kill switch.