Distributed Systems Roadmap

From what actually changes when a call crosses a machine boundary, through replication, consistency, time, consensus and delivery, to running a distributed system in production and knowing when not to build one at all. Ten levels, every lesson in the domain placed exactly once.

0 / 170 roadmap lessons mastered0%
  1. 1

    Remote Calls and Partial Failure

    0 / 18 mastered

    Stop treating a network call as a function call. Everything else in the domain depends on this shift.

  2. 2

    Replication and Consistency

    0 / 18 mastered

    Copies buy availability and cost you agreement. Learn to name the guarantee precisely.

  3. 3

    Time, Ordering and Conflict

    0 / 14 mastered

    Clocks disagree. Causality is what you can actually rely on — and when two writes are concurrent, someone must decide.

  4. 4

    Consensus and Coordination

    0 / 18 mastered

    Agreement under failure, what it assumes, what it costs, and how often you can avoid needing it.

  5. 5

    Messaging and Streams

    0 / 16 mastered

    Brokers, queues and logs — and the delivery guarantee each one actually provides.

  6. 6

    Partitioning and Membership

    0 / 14 mastered

    Splitting data across nodes, and knowing which nodes are still there.

  7. 7

    Workflows and Idempotency

    0 / 13 mastered

    Atomicity across services you do not control, and making retries safe.

  8. 8

    Overload, Deadlines and Caching

    0 / 18 mastered

    Bounded behaviour when demand exceeds capacity, and time budgets across a call graph.

    Backpressure Is a Signal That Has to Travel — and Reach Someone Who Can Slow DownOverload & BackpressureRejecting Work on Purpose — and Rejecting It Cheaply Enough to HelpOverload & BackpressureDecide at the Door Whether the Capacity ExistsOverload & BackpressureOne Retry per Tier Is Not One Retry — It MultipliesOverload & BackpressureCap Retries as a Fraction of Traffic, Not as a Count per RequestOverload & BackpressureWithout Jitter, Every Client That Failed Together Retries TogetherOverload & BackpressureContainment Is Decided by What Is Shared, Not by Where the Service Boundaries AreOverload & BackpressureBulkheads: Buying Independence by Giving Up UtilisationOverload & BackpressureA Deadline Is Divided Across the Call Chain, Not Repeated at Every HopDeadlines & Tail LatencyPass the Remaining Budget Down, Not a Fresh OneDeadlines & Tail LatencyThe Caller Is Gone — Stopping Is Usually Right and Sometimes UnsafeDeadlines & Tail LatencySend a Second Request After p95 and Take Whichever Answers FirstDeadlines & Tail LatencyFan Out to 100 and the Component’s Tail Becomes the System’s MedianDeadlines & Tail LatencyA Cache Across Machines Is a Replica With No Replication ProtocolDistributed CachingInvalidation Is a Messaging Problem, Which Is Why Cache Bugs Are HardDistributed CachingOne Key Expires and Five Hundred Instances Miss at the Same MillisecondDistributed CachingYou Cannot Enumerate the Caches, So TTL Is the Bound and Invalidation Is the OptimisationDistributed CachingSharding Does Not Help a Single KeyDistributed Caching
  9. 9

    Storage, Compute and Geography

    0 / 19 mastered

    The layers beneath a distributed database, and what physics does to a multi-region design.

  10. 10

    Operating Distributed Systems

    0 / 22 mastered

    Where to cut the system, how it fails in production, and how it recovers.

    Detect, Contain, Recover, Reconcile, VerifyFailure & Recovery in ProductionGraceful Degradation: Which Dependency Is Actually CriticalFailure & Recovery in ProductionThe Steady-State Hypothesis and the Abort ConditionFailure & Recovery in ProductionCascading Failure: When the Response to Failure Causes More FailureFailure & Recovery in ProductionDependency Blast Radius: What Dies If This Node DiesFailure & Recovery in ProductionChaos Engineering Is Not Randomly Breaking ProductionFailure & Recovery in ProductionFault Injection: The Catalogue, and Which Faults Are HardFailure & Recovery in ProductionDistributed Debugging: The Question LadderFailure & Recovery in ProductionThree Nodes, Three Logs, and You Cannot Sort by TimestampFailure & Recovery in ProductionMicroservices Are a Distribution Decision, Not a Scaling TechniqueDistribution BoundariesThe Distributed Monolith: All of the Cost, None of the AutonomyDistribution BoundariesFour Questions That Test a Proposed BoundaryDistribution BoundariesThe Shared Database: An Honest Trade, Not a ProhibitionDistribution BoundariesExactly One Component Owns Each Piece of StateDistribution BoundariesSource of Truth: The Question Every Inconsistency Incident Is Really AskingDistribution BoundariesMaterialized Views: A Read Model That LagsDistribution BoundariesReconciliation Is a Component, Not a Cleanup ScriptDistribution BoundariesAn Agent System Is a Distributed SystemAgentic Distributed SystemsThe Model Retries Because It Cannot See the ResultAgentic Distributed SystemsAgents Do Not Negotiate. Processes Contend for State.Agentic Distributed SystemsResuming a Workflow That Died Halfway ThroughAgentic Distributed SystemsThe Failures That Produce No ErrorsAgentic Distributed Systems