Cloud & Infrastructure Roadmap

Nine levels in one order. Start at Level 1, Compute, storage, networking, regions — from the workload, not a product catalogue. Every level names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.

Where to start

0 / 126 lessons masteredNot started 126Learning 0Practicing 0Mastered 0
  1. 1

    Compute, storage, networking, regions

    Start here
    0/11

    Start from the workload, not the catalogue. What an application needs before any service is named, the honest on-premises comparison, regions and failure domains, the layer stack from the application down to the data center, and where the provider's responsibility ends. Compute, storage and networking each get their fundamentals here; the depth comes in later levels.

    Before moving on: List what an application needs before naming a single service, and say where the provider's responsibility ends and yours begins.

    From Laptop to Production →
  2. 2

    VMs, containers, load balancing, DNS

    0/11

    The first thing that is not your laptop. What a hypervisor gives a guest, the lifecycle from provision to terminate, patching as an obligation, immutable images against machines you log into, and why a container shares a kernel a VM does not. Then load balancers, DNS and the CDN edge, so a request can find whichever one you chose.

    Before moving on: Explain what a VM actually is, why a container is not one, and trace a request from a DNS name through a load balancer to either.

  3. 3

    Storage types, managed databases, IAM

    0/11

    Object, block and file storage are three different contracts, and the access pattern picks between them. Managed databases come with an honest boundary: what the provider runs and what stays yours. Then identity, the deepest module in the domain: identity → policy → action → resource, human against workload credentials, policy anatomy, and least privilege stated as blast radius rather than as a checkbox. It needs the workloads from the previous level, because a workload identity has to belong to something.

    Before moving on: Pick a storage shape from the access pattern and write a workload policy narrow enough to survive a compromise.

    IAM debugging lab →
  4. 4

    Docker, registries, CI/CD

    0/10

    Packaging from the operations side, so the thing you tested is the thing that runs. Image layers and the build pipeline, why image size is a security concern, registries and digests, configuration kept out of the image and state kept out of the container. Then the pipeline itself: its identity, building an artifact once and promoting it unchanged, and the supply chain from source to the image production actually pulls.

    Before moving on: Trace an artifact from a commit to the digest production pulled, and name which identity was allowed to perform each step.

    Follow a Deployment →
  5. 5

    Infrastructure as code

    0/8

    Stop changing production by hand. Definition → plan → apply → real resources: declarative desired state against imperative scripts, Terraform's vocabulary, state as the thing that makes it work and the thing that will hurt you, drift between the file and reality, modules that help against abstraction that hides, and environments that differ on purpose. It comes after the pipeline because the pipeline is what runs apply, under the identity the previous levels gave it.

    Before moving on: Read a plan, explain what state is for, and predict what the next apply does to a resource someone created by hand.

    Drift lab →
  6. 6

    Networking boundaries, orchestration, autoscaling

    0/33

    The largest level, in three parts. Virtual networks, public and private subnets, route tables, gateways, stateful security groups against stateless ACLs and private connectivity draw the trust boundaries. Orchestration is then derived from the problem — placement, restarts, health, rollout, discovery — with Kubernetes as a control loop over desired state, scheduling with requests and limits, and the mandatory lesson on when none of it is warranted. Serverless and autoscaling close it: the signal that represents pressure, startup time as the reason capacity is always late, and the liveness/readiness confusion that turns a deploy into an outage. It needs the images from level 4, because an orchestrator schedules images, not code.

    Before moving on: Justify or refuse orchestration for a given workload, place it in the right subnet, and explain why new capacity is always late.

    Do we need Kubernetes? →
  7. 7

    Reliability, multi-zone, backup and recovery

    0/14

    Redundancy that is actually redundant. Failure domains from process to region, multi-zone topologies, RPO and RTO as requirements rather than adjectives, backup strategy and the restore you have never tested. Deployment strategies sit here too: rolling, blue-green and canary are reliability tools, each depends on shutting down without dropping in-flight work, and each assumes the promoted artifact and the health checks of the earlier levels.

    Before moving on: Find the failure domain hiding behind three replicas, pick a rollout strategy for a schema change, and restore a backup on purpose.

    Break This Infrastructure →
  8. 8

    Multi-region, cloud security, cost engineering

    0/19

    The expensive decisions, with their costs stated out loud. Multi-region and the data problem active-active creates; the security overlay every diagram gets — public exposure with context, trust boundaries, roles instead of static keys, secrets and the key hierarchy behind encryption at rest; the infrastructure signals that say the platform under the application is healthy; and cost as a design constraint, from idle capacity and right-sizing to egress and storage that ages into a cheaper tier.

    Before moving on: Present an architecture in four views — structure, reliability, security and cost — and defend each one.

    Cost engineering →
  9. 9

    Production infrastructure design

    0/9

    Receive "we built an application, how do we run it?" and answer it systematically. What is different about model and agent workloads, then the strategy questions: multi-cloud and hybrid taught with their real cost, migration as inventory → dependency mapping → strategy → pilot → gradual move → validation, and an operational-complexity score that makes "we added Kubernetes and a service mesh" visible as a decision. The capstone at the end asks the question the whole domain exists to answer.

    Before moving on: Design infrastructure that meets the requirement and argue convincingly for everything you chose not to build.

    Production SaaS capstone →