Cloud & Infrastructure Roadmap
Nine levels in one order. Start at Level 1, Compute, storage, networking, regions — from the workload, not a product catalogue. Every level names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.
Where to start
Cloud & Infrastructure
9 stages · 0/126 lessonsCompute, networking, storage, identity, delivery, reliability and cost — and the judgement to refuse complexity the workload never asked for.
- Compute, storage, networking, regions
- VMs, containers, load balancing, DNS
- Storage types, managed databases, IAM
- Docker, registries, CI/CD
- Infrastructure as code
- Networking boundaries, orchestration, autoscaling
- Reliability, multi-zone, backup and recovery
- Multi-region, cloud security, cost engineering
- Production infrastructure design
- 10/11
Compute, storage, networking, regions
Start hereStart from the workload, not the catalogue. What an application needs before any service is named, the honest on-premises comparison, regions and failure domains, the layer stack from the application down to the data center, and where the provider's responsibility ends. Compute, storage and networking each get their fundamentals here; the depth comes in later levels.
Before moving on: List what an application needs before naming a single service, and say where the provider's responsibility ends and yours begins.
From Laptop to Production → - 20/11
VMs, containers, load balancing, DNS
The first thing that is not your laptop. What a hypervisor gives a guest, the lifecycle from provision to terminate, patching as an obligation, immutable images against machines you log into, and why a container shares a kernel a VM does not. Then load balancers, DNS and the CDN edge, so a request can find whichever one you chose.
Before moving on: Explain what a VM actually is, why a container is not one, and trace a request from a DNS name through a load balancer to either.
Needs first:Compute, storage, networking, regions - 30/11
Storage types, managed databases, IAM
Object, block and file storage are three different contracts, and the access pattern picks between them. Managed databases come with an honest boundary: what the provider runs and what stays yours. Then identity, the deepest module in the domain: identity → policy → action → resource, human against workload credentials, policy anatomy, and least privilege stated as blast radius rather than as a checkbox. It needs the workloads from the previous level, because a workload identity has to belong to something.
Before moving on: Pick a storage shape from the access pattern and write a workload policy narrow enough to survive a compromise.
IAM debugging lab → - 40/10
Docker, registries, CI/CD
Packaging from the operations side, so the thing you tested is the thing that runs. Image layers and the build pipeline, why image size is a security concern, registries and digests, configuration kept out of the image and state kept out of the container. Then the pipeline itself: its identity, building an artifact once and promoting it unchanged, and the supply chain from source to the image production actually pulls.
Before moving on: Trace an artifact from a commit to the digest production pulled, and name which identity was allowed to perform each step.
Follow a Deployment →- What Is Inside a Container Image
- Docker Fundamentals — One Implementation of the Model
- The Container Build Pipeline
- Why Image Size Is an Infrastructure Problem
- The Container Registry
- Configuration Belongs Outside the Image
- Persistent Data and Containers
- The Pipeline as Infrastructure
- Build Once, Promote the Same Bytes
- The Infrastructure Supply Chain
- 50/8
Infrastructure as code
Stop changing production by hand. Definition → plan → apply → real resources: declarative desired state against imperative scripts, Terraform's vocabulary, state as the thing that makes it work and the thing that will hurt you, drift between the file and reality, modules that help against abstraction that hides, and environments that differ on purpose. It comes after the pipeline because the pipeline is what runs
apply, under the identity the previous levels gave it.Before moving on: Read a plan, explain what state is for, and predict what the next apply does to a resource someone created by hand.
Drift lab →- Infrastructure as Code
- Declarative vs Imperative Infrastructure
- Terraform: The Vocabulary of Declarative Infrastructure
- State: The File That Makes It Work and the File That Will Hurt You
- Reading a Plan Before You Apply It
- Drift: When the File and Reality Disagree
- Modules: Reuse Without Hiding
- Development, Staging and Production
- 60/33
Networking boundaries, orchestration, autoscaling
The largest level, in three parts. Virtual networks, public and private subnets, route tables, gateways, stateful security groups against stateless ACLs and private connectivity draw the trust boundaries. Orchestration is then derived from the problem — placement, restarts, health, rollout, discovery — with Kubernetes as a control loop over desired state, scheduling with requests and limits, and the mandatory lesson on when none of it is warranted. Serverless and autoscaling close it: the signal that represents pressure, startup time as the reason capacity is always late, and the liveness/readiness confusion that turns a deploy into an outage. It needs the images from level 4, because an orchestrator schedules images, not code.
Before moving on: Justify or refuse orchestration for a given workload, place it in the right subnet, and explain why new capacity is always late.
Do we need Kubernetes? →- Virtual Private Cloud
- Public and Private Subnets
- Route Tables
- Internet Gateway
- NAT Gateway
- Security Groups: The Stateful Firewall
- Network ACLs: The Stateless Filter
- Private Connectivity
- Why Orchestration Exists
- Kubernetes: Why It Exists
- The Kubernetes Mental Model
- Core Objects, and Why Each One Exists
- Pods: Shared Lifecycle, Shared Network
- Deployments and the Replica Controller
- Self-Healing, and What It Does Not Heal
- Service: A Stable Name in Front of Moving Pods
- Ingress and Gateway: Getting Traffic In
- ConfigMap vs Secret — and the Honest Limit of a Secret
- Stateful Workloads: Databases Are Not Stateless APIs
- Scheduling: How a Pod Chooses a Node
- Requests vs Limits: Two Numbers That Do Different Jobs
- OOM Kills and CPU Throttling
- Horizontal Pod Autoscaling — and Why New Capacity Is Always Late
- Kubernetes Is Not Always Needed
- Serverless as an Execution Model
- Serverless Trade-offs
- Serverless and Database Connections
- Choosing a Compute Model
- Autoscaling
- Autoscaling Signals
- Startup Time & Cold Start
- Health Checks
- Liveness vs Readiness
- 70/14
Reliability, multi-zone, backup and recovery
Redundancy that is actually redundant. Failure domains from process to region, multi-zone topologies, RPO and RTO as requirements rather than adjectives, backup strategy and the restore you have never tested. Deployment strategies sit here too: rolling, blue-green and canary are reliability tools, each depends on shutting down without dropping in-flight work, and each assumes the promoted artifact and the health checks of the earlier levels.
Before moving on: Find the failure domain hiding behind three replicas, pick a rollout strategy for a schema change, and restore a backup on purpose.
Break This Infrastructure →- Infrastructure Reliability
- High Availability
- Failure Domains
- Multi-Zone Deployment
- Disaster Recovery
- RPO & RTO
- Backup Strategy
- Restore Testing
- Graceful Shutdown: The 502 Spike Nobody Investigates
- Four Ways to Replace Running Code
- Rolling Deployment and the Compatibility It Demands
- Blue/Green: Two Environments, One Switch
- Canary: Let 5% of Traffic Find the Bug
- Deployment Is Not Release
- 80/19
Multi-region, cloud security, cost engineering
The expensive decisions, with their costs stated out loud. Multi-region and the data problem active-active creates; the security overlay every diagram gets — public exposure with context, trust boundaries, roles instead of static keys, secrets and the key hierarchy behind encryption at rest; the infrastructure signals that say the platform under the application is healthy; and cost as a design constraint, from idle capacity and right-sizing to egress and storage that ages into a cheaper tier.
Before moving on: Present an architecture in four views — structure, reliability, security and cost — and defend each one.
Needs first:Storage types, managed databases, IAMNetworking boundaries, orchestration, autoscalingReliability, multi-zone, backup and recoveryCost engineering →- Multi-Region Deployment
- Active-Passive Failover
- Active-Active
- The Security View
- Public Exposure, Read With Context
- Infrastructure Trust Boundaries
- Roles vs Static Keys
- Secrets in Infrastructure
- Key Management and Encryption at Rest
- Infrastructure Observability
- Infrastructure Logs
- Audit Trails
- Cost Engineering
- Fixed vs Variable Cost
- Idle Capacity: Headroom or Waste?
- Right-Sizing Without Causing an Outage
- Cost per Service and the Attribution Problem
- Egress: Moving Data Costs Money, Not Just Storing It
- Storage Lifecycle: Hot, Warm, Archive, Delete
- 90/9
Production infrastructure design
Receive "we built an application, how do we run it?" and answer it systematically. What is different about model and agent workloads, then the strategy questions: multi-cloud and hybrid taught with their real cost, migration as inventory → dependency mapping → strategy → pilot → gradual move → validation, and an operational-complexity score that makes "we added Kubernetes and a service mesh" visible as a decision. The capstone at the end asks the question the whole domain exists to answer.
Before moving on: Design infrastructure that meets the requirement and argue convincingly for everything you chose not to build.
Needs first:Multi-region, cloud security, cost engineeringProduction SaaS capstone →