Concurrency & Parallelism Roadmap
Start at Concurrency, parallelism, tasks and threads and follow the nine stages in order, ending with production debugging. Every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.
Where to start
Concurrency & Parallelism
9 stages · 0/154 lessonsOverlapping and simultaneous execution: what is shared, what must stay true, which primitive protects it, and when to leave the work sequential.
- Concurrency, parallelism, tasks and threads
- Shared state, races and critical sections
- Mutexes, semaphores and condition variables
- Async, promises, coroutines and event loops
- Thread pools, worker pools, producer/consumer and backpressure
- Deadlocks, starvation, livelock and contention
- Atomics, memory ordering and lock-free
- Parallel algorithms, fork/join and SIMD
- Parallel performance, debugging and production concurrency
- 10/16
Concurrency, parallelism, tasks and threads
Start hereThe distinction the whole domain rests on: overlapping progress versus simultaneous execution, why each exists, and how to classify work as waiting-bound or compute-bound. Then the units a runtime hands you — processes, threads, tasks and coroutines — and what each one buys in CPython, JavaScript and C++. Everything later assumes you pick the execution model from the shape of the work rather than from habit.
Before moving on: Look at a piece of work, say whether it is waiting-bound or compute-bound, and name the execution model that follows — including the case where the answer is to run it sequentially.
The scheduler lab →- What Concurrency Actually Is
- Which One Does This Workload Need?
- Why Concurrency Exists: Waiting
- Why Parallelism Exists: Compute
- Classifying the Work: Computing or Waiting?
- Choosing an Execution Model
- What "Making Progress" Actually Means
- Concurrency Is Always Bought With Complexity
- Processes: Isolation You Cannot Accidentally Break
- Threads: One Address Space, Several Instruction Streams
- Thread or Process?
- A Task Is Not a Thread
- Coroutines: Functions That Can Pause
- Python: Threads, Processes and the GIL
- JavaScript: One Event Loop per Agent, Not One Thread per Runtime
- C++: Threads, Atomics and a Memory Model With Teeth
- 20/15
Shared state, races and critical sections
Shared mutable state is where concurrency goes wrong, so it comes before any primitive. Write down the invariant, find the smallest critical section, enumerate interleavings until one breaks the invariant, and tell a race condition from a data race. Immutability, copying and optimistic control sit here because removing the sharing is often the better fix than protecting it.
Before moving on: Given a function that two threads call, write down what must stay true, produce an interleaving that violates it, and say whether the fix is a lock or the removal of the sharing.
Needs first:Concurrency, parallelism, tasks and threadsBreak the counter →- Shared Mutable State
- Interleavings: The Schedule Is Part of the Program
- Invariants: Name It Before You Lock It
- Finding the Critical Section
- Reasoning About Races: A Method, Not an Instinct
- Data Race Is Not Race Condition
- The Atomicity Illusion
- Nondeterminism: Same Input, Different Output
- Immutability as a Concurrency Strategy
- Copy or Share?
- Copy-on-Write as a Concurrency Strategy
- Optimistic Concurrency Control
- Pessimistic Concurrency
- Optimistic vs Pessimistic
- Initialization Races
- 30/11
Mutexes, semaphores and condition variables
The primitives, chosen by the guarantee you need and held over the smallest region that preserves the invariant: mutexes and lock scope, read/write locks, semaphores as permit counters, and condition variables with the predicate loop that makes lost and spurious wakeups harmless. Spin locks, reentrancy and what "thread-safe" actually claims close the stage.
Before moving on: State what a mutex guarantees and what it does not, explain why a condition variable is always waited on in a loop, and say which operations a "thread-safe" type actually protects.
Needs first:Shared state, races and critical sectionsFix the counter →- Mutexes: What They Protect and What They Do Not
- Lock Scope: What You Hold It Across
- Read/Write Locks, Honestly
- Semaphores: Counting Permits as a Resource Limit
- Semaphore versus Mutex: Not the Same Primitive
- Condition Variables: Waiting Until a Predicate Is True
- Lost Wakeups: The Notify That Arrived Before the Wait
- Spurious Wakeups: Why It Is `while`, Not `if`
- Spin Locks
- What "Thread-Safe" Actually Means
- Reentrancy
- 40/16
Async, promises, coroutines and event loops
Async/await as an execution model rather than syntax: what suspending a task actually does, what keeps running while it is suspended, and why async is not parallelism. Event loops, promises, worker threads and web workers come first; structured concurrency, cancellation, timeouts and deadlines then answer who owns a task once it is started. It only needs the execution models of the first stage.
Before moving on: Explain why an
awaitin a loop is a latency bug and an unboundedPromise.allis a capacity bug, and why a task with no owner is the async equivalent of a leaked thread.Needs first:Concurrency, parallelism, tasks and threadsStall the event loop →- Event Loops as a Concurrency Model
- Await Is a Yield Point
- Futures & Promises
- Worker Threads
- Web Workers
- Event Loop or Threads?
- Async Is Not Parallelism
- Blocking the Event Loop
- Promise.all & gather
- The Sequential Await Trap
- Structured Concurrency
- Cancellation
- Cancellation Propagation
- Timeouts
- Deadlines vs Timeouts
- Orphaned Tasks
- 50/17
Thread pools, worker pools, producer/consumer and backpressure
Concurrency as a budget you set rather than a number that emerges. Producer/consumer is the flagship pattern: bounded versus unbounded queues, backpressure, channels, message passing and actors move data instead of sharing it, which is why this follows the primitives. Thread and worker pools, sizing, work stealing and saturation decide how many workers run and what happens to the work that does not fit, for threads and for async tasks alike.
Before moving on: Size a pool by reasoning rather than by folklore, and say exactly what a full queue should do — block, drop or reject — and defend the choice.
Set a concurrency budget → - 60/15
Deadlocks, starvation, livelock and contention
How a system full of correct code stops making progress, and how to make the failure structurally impossible rather than unlikely. The four deadlock conditions and the one lock ordering removes, wait-for graphs, livelock, starvation, fairness and priority inversion; then the contention side — convoys, false parallelism, oversubscription, context-switch cost and busy waiting — which explains why adding threads made it slower. It needs the locks of the primitives stage and the pools of the previous one.
Before moving on: Draw a wait-for graph from a stack trace, name which of the four deadlock conditions your fix removes, and explain why the eight-core box was slower than the two-core one.
Needs first:Mutexes, semaphores and condition variablesThread pools, worker pools, producer/consumer and backpressureDeadlock lab → - 70/15
Atomics, memory ordering and lock-free
What an atomic operation actually makes indivisible, and why another thread might not see your write at all. Compare-and-swap, a lock-free stack, the ABA problem and wait-free versus lock-free as progress guarantees rather than speed claims; then the memory model — happens-before, reordering, barriers, false sharing, cache coherence and safe publication — with double-checked locking as the cautionary tale. It builds on the invariants of the shared-state stage and the primitives it tries to replace.
Before moving on: Explain why
Watch a CAS loop retry →counter.fetch_add(1)is safe andif (!cache) cache = build()on an atomic pointer is not, and why the fix for double-checked locking is a memory-model question rather than a syntax one.- Atomics: What Is Actually Indivisible
- Compare-and-Swap and the Retry Loop
- Atomics Are Not Magic
- Lock-Free Is a Progress Guarantee
- A Lock-Free Stack, and What the Teaching Version Omits
- Wait-Free vs Lock-Free: Whose Progress Is Guaranteed
- The ABA Problem: The Value Came Back
- What a Memory Model Defines
- Happens-Before: The Edge That Makes a Write Visible
- Reordering: The Compiler and the CPU Both Do It
- Memory Barriers Constrain Ordering, Not Caches
- False Sharing: Different Variables, Same Cache Line
- What a Shared Write Costs
- Safe Publication: Handing Over a Finished Object
- Double-Checked Locking: The Canonical Cautionary Tale
- 80/18
Parallel algorithms, fork/join and SIMD
Decompose a computation into independent work and predict the ceiling before writing any of it. Fork/join, parallel reduce and map/reduce; the overhead that makes small work slower in parallel; Amdahl, Gustafson and work versus span as the limits; SIMD, pipeline parallelism, fan-out/fan-in and scatter/gather; barriers and latches for agreeing on when to proceed. The workers come from the pools stage, and the contention stage explains why the speedup stops early.
Before moving on: Compute the best speedup a decomposition can ever reach from its serial fraction and its span, and say when the parallel version will lose to the sequential one.
Needs first:Thread pools, worker pools, producer/consumer and backpressureDeadlocks, starvation, livelock and contentionScaling lab →- Fork/Join
- Parallel Reduce
- The Map/Reduce Pattern
- Parallel Algorithms
- Parallel Overhead
- Amdahl's Law
- Gustafson's Law
- Work and Span
- Dependency Graphs
- SIMD: One Instruction, Many Elements
- Task Parallelism vs Data Parallelism
- GPU Parallelism: Thousands of Lanes, One Bus
- Pipeline Parallelism: Different Items, Different Stages
- Fan-Out / Fan-In: One Request Becomes N
- Scatter/Gather and the Tail You Inherit
- Barriers
- Latches & Countdowns
- Parallelism Moves the Load Downstream
- 90/31
Parallel performance, debugging and production concurrency
The production stage: "it is slow and occasionally wrong, and it only happens in production", answered with evidence instead of guesses. Why scaling is not linear — memory bandwidth, cache locality, NUMA and affinity; observability, lock-wait metrics, thread and task dumps, profilers, race detectors and deterministic replay for the heisenbugs; then the server, UI, database, distributed and agent settings where a concurrency model has to be chosen and defended. It draws on every earlier stage, the failure, memory-model and parallel ones most.
Before moving on: Instrument a concurrent system so the next failure leaves evidence, and design one whose concurrency model you can defend line by line — including the parts you chose to leave sequential.
Needs first:Deadlocks, starvation, livelock and contentionAtomics, memory ordering and lock-freeParallel algorithms, fork/join and SIMDConcurrency capstone →- Why Eight Cores Give You Four and a Half
- Memory Bandwidth: More Cores, Same Bus
- Parallelism Can Destroy Locality
- NUMA: Not All Memory Costs the Same
- Thread Affinity: Pinning, and What It Costs You
- Reduction Ordering: The Sum Changed When the Worker Count Did
- Determinism: Same Input, Same Output?
- Ordering Guarantees: Four Levels, Four Prices
- What to Instrument in a Concurrent System
- Hold Time, Wait Time, and the Ratio Between Them
- Reading a Thread Dump
- Task Dumps: When the Threads Look Idle and Nothing Is Moving
- Off-CPU Time: The Thing a CPU Profiler Cannot See
- Race Detectors: What They Find, and What They Structurally Cannot
- Heisenbugs: The Bug That Leaves When You Look at It
- Deterministic Replay: Making the Schedule Reproducible
- Stress Testing: A Test That Passed Once Proves Nothing
- Choosing a Concurrency Model for a Server
- Thread per Request: The Model That Reads Like Ordinary Code
- Event-Driven Servers: Many Connections, One Loop
- Hybrid Runtimes: It Was Never Threads Versus Async
- UI Concurrency: One Thread Owns the Screen
- The Database Solves Concurrency For Its Data, Not For Your Memory
- What Changes When the Shared State Is on Another Machine
- A Mutex on Server A Does Nothing About Server B
- The Concurrency Pattern Catalogue
- Concurrency Anti-Patterns
- Can These Tool Calls Run At the Same Time?
- Parallel Tool Calls: Lower Latency, and Four Bills
- Two Agents, One Document
- Cancelling an Agent Run: What Actually Stops