Concurrency & Parallelism Roadmap

Start at Concurrency, parallelism, tasks and threads and follow the nine stages in order, ending with production debugging. Every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.

Where to start

0 / 154 lessons masteredNot started 154Learning 0Practicing 0Mastered 0
  1. 1

    Concurrency, parallelism, tasks and threads

    Start here
    0/16

    The distinction the whole domain rests on: overlapping progress versus simultaneous execution, why each exists, and how to classify work as waiting-bound or compute-bound. Then the units a runtime hands you — processes, threads, tasks and coroutines — and what each one buys in CPython, JavaScript and C++. Everything later assumes you pick the execution model from the shape of the work rather than from habit.

    Before moving on: Look at a piece of work, say whether it is waiting-bound or compute-bound, and name the execution model that follows — including the case where the answer is to run it sequentially.

    The scheduler lab →
  2. 2

    Shared state, races and critical sections

    0/15

    Shared mutable state is where concurrency goes wrong, so it comes before any primitive. Write down the invariant, find the smallest critical section, enumerate interleavings until one breaks the invariant, and tell a race condition from a data race. Immutability, copying and optimistic control sit here because removing the sharing is often the better fix than protecting it.

    Before moving on: Given a function that two threads call, write down what must stay true, produce an interleaving that violates it, and say whether the fix is a lock or the removal of the sharing.

    Break the counter →
  3. 3

    Mutexes, semaphores and condition variables

    0/11

    The primitives, chosen by the guarantee you need and held over the smallest region that preserves the invariant: mutexes and lock scope, read/write locks, semaphores as permit counters, and condition variables with the predicate loop that makes lost and spurious wakeups harmless. Spin locks, reentrancy and what "thread-safe" actually claims close the stage.

    Before moving on: State what a mutex guarantees and what it does not, explain why a condition variable is always waited on in a loop, and say which operations a "thread-safe" type actually protects.

    Fix the counter →
  4. 4

    Async, promises, coroutines and event loops

    0/16

    Async/await as an execution model rather than syntax: what suspending a task actually does, what keeps running while it is suspended, and why async is not parallelism. Event loops, promises, worker threads and web workers come first; structured concurrency, cancellation, timeouts and deadlines then answer who owns a task once it is started. It only needs the execution models of the first stage.

    Before moving on: Explain why an await in a loop is a latency bug and an unbounded Promise.all is a capacity bug, and why a task with no owner is the async equivalent of a leaked thread.

    Stall the event loop →
  5. 5

    Thread pools, worker pools, producer/consumer and backpressure

    0/17

    Concurrency as a budget you set rather than a number that emerges. Producer/consumer is the flagship pattern: bounded versus unbounded queues, backpressure, channels, message passing and actors move data instead of sharing it, which is why this follows the primitives. Thread and worker pools, sizing, work stealing and saturation decide how many workers run and what happens to the work that does not fit, for threads and for async tasks alike.

    Before moving on: Size a pool by reasoning rather than by folklore, and say exactly what a full queue should do — block, drop or reject — and defend the choice.

    Set a concurrency budget →
  6. 6

    Deadlocks, starvation, livelock and contention

    0/15

    How a system full of correct code stops making progress, and how to make the failure structurally impossible rather than unlikely. The four deadlock conditions and the one lock ordering removes, wait-for graphs, livelock, starvation, fairness and priority inversion; then the contention side — convoys, false parallelism, oversubscription, context-switch cost and busy waiting — which explains why adding threads made it slower. It needs the locks of the primitives stage and the pools of the previous one.

    Before moving on: Draw a wait-for graph from a stack trace, name which of the four deadlock conditions your fix removes, and explain why the eight-core box was slower than the two-core one.

    Deadlock lab →
  7. 7

    Atomics, memory ordering and lock-free

    0/15

    What an atomic operation actually makes indivisible, and why another thread might not see your write at all. Compare-and-swap, a lock-free stack, the ABA problem and wait-free versus lock-free as progress guarantees rather than speed claims; then the memory model — happens-before, reordering, barriers, false sharing, cache coherence and safe publication — with double-checked locking as the cautionary tale. It builds on the invariants of the shared-state stage and the primitives it tries to replace.

    Before moving on: Explain why counter.fetch_add(1) is safe and if (!cache) cache = build() on an atomic pointer is not, and why the fix for double-checked locking is a memory-model question rather than a syntax one.

    Watch a CAS loop retry →
  8. 8

    Parallel algorithms, fork/join and SIMD

    0/18

    Decompose a computation into independent work and predict the ceiling before writing any of it. Fork/join, parallel reduce and map/reduce; the overhead that makes small work slower in parallel; Amdahl, Gustafson and work versus span as the limits; SIMD, pipeline parallelism, fan-out/fan-in and scatter/gather; barriers and latches for agreeing on when to proceed. The workers come from the pools stage, and the contention stage explains why the speedup stops early.

    Before moving on: Compute the best speedup a decomposition can ever reach from its serial fraction and its span, and say when the parallel version will lose to the sequential one.

    Scaling lab →
  9. 9

    Parallel performance, debugging and production concurrency

    0/31

    The production stage: "it is slow and occasionally wrong, and it only happens in production", answered with evidence instead of guesses. Why scaling is not linear — memory bandwidth, cache locality, NUMA and affinity; observability, lock-wait metrics, thread and task dumps, profilers, race detectors and deterministic replay for the heisenbugs; then the server, UI, database, distributed and agent settings where a concurrency model has to be chosen and defended. It draws on every earlier stage, the failure, memory-model and parallel ones most.

    Before moving on: Instrument a concurrent system so the next failure leaves evidence, and design one whose concurrency model you can defend line by line — including the parts you chose to leave sequential.

    Concurrency capstone →