Software Architecture Roadmap

Start with the two prerequisites taught elsewhere, then APIs — the first stage with its own lessons. Every stage names what it needs first and what you should be able to do before moving on, and stages taught in another domain link there. Progress is stored locally in your browser.

Where to start

Software architecture

17 stages · 0/37 lessons

How to structure a software system and why: from one process to a fleet of services, adding a component only when a real problem asks for it.

  1. Software Fundamentals
  2. HTTP / Networking
  3. APIs
  4. Application Architecture
  5. Databases
  6. Caching
  7. Queues
  8. Async Processing
  9. Scaling
  10. Load Balancing
  11. Distributed Systems
  12. Microservices
  13. Events
  14. Reliability
  15. Observability
  16. System Design
  17. Production Architecture
0 / 37 lessons masteredNot started 37Learning 0Practicing 0Mastered 0
  1. 1

    Software Fundamentals

    Start hereTaught elsewhere

    Data structures and algorithms: hash tables, queues, graphs, trees — the ideas that reappear later as consistent hashing, message queues, dependency graphs and indexes. Taught in the DSA domain; come here first so those mechanisms feel familiar when they show up at system scale.

    Before moving on: Say what a hash table, a queue, a tree and a graph cost per operation, and recognise each one when it reappears as a piece of infrastructure.

    Data Structures & Algorithms →
  2. 2

    HTTP / Networking

    Taught elsewhere

    What to know before the rest makes sense: TCP connections and their cost, HTTP/1.1 vs HTTP/2 multiplexing, TLS termination, DNS and TTLs, status codes, the headers that control caching (Cache-Control, ETag), timeouts at every hop, and why a timeout is indistinguishable from a partition. Taught in the Networking domain.

    Before moving on: Explain what a TCP connection and a TLS handshake cost, read Cache-Control and ETag headers, and say why a timeout looks the same as a network partition.

    Computer Networking roadmap →
  3. 3

    APIs

    0/1

    REST, GraphQL, gRPC, WebSockets, SSE and webhooks: what each solves, how each fails, and which to expose to whom. The first stage with lessons of its own — an API is the boundary every later component sits behind, so learn the boundary before the components.

    Before moving on: Pick REST, GraphQL, gRPC, WebSockets, SSE or webhooks for a given consumer and say how each one fails on the wire and under load.

    Needs first:HTTP / Networking
  4. 4

    Application Architecture

    0/8

    What architecture is, how a system evolves one problem at a time, monolith and modular monolith, and the code-level structures — layered, clean, hexagonal — that keep dependencies pointing the right way. This is the single-process baseline; every stage after this one adds a component to it and pays for it.

    Before moving on: Draw the components and dependency direction of a system, defend a monolith or modular monolith for a small team, and place a new feature in the right layer.

    Needs first:APIs
  5. 5

    Databases

    Taught elsewhere

    Indexes, transactions, replication and partitioning: the database is the ceiling of most systems, and its behaviour drives the architecture around it. Taught in the Database Engineering domain; it comes right after the application because caching, scaling and distributed data are all reactions to what the database can and cannot do.

    Before moving on: Choose an index for a query, explain what a transaction guarantees, and say what replication lag and a partition key do to the code around the database.

    Database Engineering roadmap →
  6. 6

    Caching

    0/2

    Browser → CDN → application cache → Redis → database: TTLs, invalidation, stampedes, hot keys, and what belongs at the edge. The first thing you add to protect the database, and it reuses the HTTP caching headers from the networking stage.

    Before moving on: Place a cache at the right layer, choose a TTL and an invalidation strategy, and prevent a stampede on a hot key.

  7. 7

    Queues

    0/2

    Point-to-point queues with acks, retries and dead letters, then Kafka-style logs with partitions, offsets and consumer groups. The queue from DSA becomes a component between two parts of the application, and it is the mechanism async processing and events are built on.

    Before moving on: Design a producer and a consumer with acks, retries and a dead-letter queue, and explain what a Kafka partition and a consumer group guarantee about ordering.

  8. 8

    Async Processing

    0/3

    Background jobs and workers, idempotency under retries, and backpressure when producers outrun consumers. Once work goes through a queue it will be retried, so this stage is where "safe to run twice" becomes a design requirement.

    Before moving on: Move slow work to a worker, make the job safe to run twice, and stop a backlog from taking the system down.

    Needs first:Queues
  9. 9

    Scaling

    0/3

    Vertical vs horizontal, why stateless services scale and stateful ones do not, and scaling a system step by step from one server to many. It needs the cache and the database stages first, because the walk-through adds those components in the order a real system does.

    Before moving on: Take a system from one server to many one bottleneck at a time, and say which state has to leave the service before it can scale out.

    Needs first:DatabasesCaching
  10. 10

    Load Balancing

    0/1

    Round robin, least connections and consistent hashing; health checks, sticky sessions, L4 vs L7, and the balancer as a single point of failure. Comes right after scaling: many stateless instances need something in front of them.

    Before moving on: Choose an algorithm and a layer (L4 or L7) for a workload, configure health checks, and remove the balancer as a single point of failure.

    Needs first:Scaling
  11. 11

    Distributed Systems

    0/4

    Partitions and CAP without slogans, consistent hashing, service discovery in a dynamic fleet, and why one transaction cannot span services. Load balancing gave you a fleet; the database stage gave you transactions; this stage is what happens to both when the network is unreliable.

    Before moving on: Explain what a partition forces you to give up, hash keys onto a fleet that changes size, find services without hard-coded addresses, and argue why a transaction cannot span two services.

  12. 12

    Microservices

    0/2

    Service boundaries, database ownership, partial failure and the API gateway — and the distributed monolith that results when the boundaries are not real. Deliberately late: it is the monolith from the application stage split along boundaries, with every distributed-systems cost attached.

    Before moving on: Draw service boundaries that own their data, say what belongs in the gateway, and spot a distributed monolith from its deployment story.

  13. 13

    Events

    0/5

    Event-driven architecture, the sync-vs-async decision, sagas with compensation, CQRS, and event sourcing — with duplicates, ordering and replay taken seriously. Built on the queue stage, and needed once services own separate databases and still have to agree.

    Before moving on: Choose sync or async per interaction, write a saga with compensation, and handle duplicate, out-of-order and replayed events.

  14. 14

    Reliability

    0/4

    Timeouts, retries with budgets, circuit breakers, bulkheads, rate limiting, and the arithmetic of nines and error budgets. Every hop added since the application stage is a place to fail; this stage is the toolkit for failing on purpose instead of by accident, and it assumes the idempotency work from async processing.

    Before moving on: Set timeouts and retry budgets per hop, wire a circuit breaker and a rate limiter, and turn an availability target into an error budget.

  15. 15

    Observability

    0/2

    Logs, metrics and traces for the same incident, and following one request across five hops to find where the time went. Only makes sense once there are hops to follow, so it comes after microservices.

    Before moving on: Correlate a log line, a metric and a trace for one incident, and read a distributed trace to find the slow hop.

    Needs first:Microservices
  16. 16

    System Design

    Taught elsewhere

    Put it together: URL shortener, chat, social feed, e-commerce, notifications, file storage — requirements in, an evolving architecture out. The exercises use every stage above, so they come last.

    Before moving on: Turn a requirements brief into an architecture with estimates, an API, a data model and its failure modes named, adding one component per problem.

    System design exercises →
  17. 17

    Production Architecture

    Taught elsewhere

    Kill components and watch what breaks: the failure simulator turns every pattern above into an on-call decision. It closes the roadmap because reading its output takes the reliability and observability vocabulary.

    Before moving on: Predict which user flows break when a component dies, and name the mitigation that keeps each flow degraded instead of failed.

    Failure simulator →