Computer Architecture Roadmap
Start at Level 1 and follow the nine levels in order, from building arithmetic out of logic gates to reading performance counters on a loaded multicore machine. Every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.
Where to start
Computer Architecture
9 stages · 0/138 lessonsHow software actually executes on real hardware, and why data movement so often matters more than arithmetic.
- Level 1 · How arithmetic becomes hardware
- Level 2 · What a CPU is made of
- Level 3 · The contract, and executing against it
- Level 4 · Guessing, reordering and doing several things at once
- Level 5 · The memory hierarchy
- Level 6 · Past the cache, and where the bytes sit
- Level 7 · Address translation and protection
- Level 8 · More than one core
- Level 9 · The rest of the machine, and reading it
- 10/11
Level 1 · How arithmetic becomes hardware
Start hereBinary and two's complement, why floating point approximates, and the path from a truth table to an adder. It comes first because every later level treats a register, an ALU and a memory cell as given; here you see they are made of gates. By the end, arithmetic is no longer a primitive — it is something you could build.
Before moving on: Convert between decimal and two's complement by hand, explain why
0.1 + 0.2is not0.3, and draw a one-bit full adder out of gates.- What Actually Happens When You Add Two Numbers
- Binary, and Why Everything Is Eventually Bits
- Bits, Bytes and Words — and Why "Word" Is Not a Fixed Size
- Two's Complement: One Circuit for Addition and Subtraction
- Integer Overflow: The Hardware Wraps, the Language Decides
- Floating Point: Trading Precision for Range
- Why 0.1 + 0.2 Is Not 0.2 + 0.1's Problem
- Boolean Logic: The Four Operations Hardware Actually Has
- Logic Gates: Where Software Stops and Physics Starts
- Building an Adder: Where Arithmetic Comes From
- Adding Memory: Combinational, Sequential and the Clock
- 20/7
Level 2 · What a CPU is made of
Registers, the ALU, the datapath they sit on and the control unit that steers it, wired together from the adders and latches of Level 1. Plus the first myth to lose: that clock speed is performance.
Before moving on: Name what each part of a CPU does while one instruction runs, and explain why two CPUs at the same clock can finish the same program in different times.
Needs first:Level 1 · How arithmetic becomes hardware- What Is Actually Inside a CPU
- Registers: The Fastest Storage, and There Is Almost None of It
- The ALU: Where Arithmetic Actually Happens
- The Control Unit: Turning Instructions Into Actions
- The Datapath: How Values Move Through the Machine
- The Program Counter: Deciding What Happens Next
- The Clock: Why GHz Is Not Performance
- 30/14
Level 3 · The contract, and executing against it
The ISA as the interface between software and hardware, the distinction from microarchitecture that most performance arguments miss, and the fetch-decode-execute cycle that pipelining then overlaps. It needs the datapath from Level 2, because a pipeline is that datapath cut into stages.
Before moving on: Read a short assembly listing, say what the ISA fixes and what a microarchitecture is free to change, and trace an instruction through a five-stage pipeline including a stall and a forward.
Needs first:Level 2 · What a CPU is made of- The ISA: The Contract Between Software and Hardware
- ISA vs Microarchitecture: The Distinction Everything Depends On
- x86-64, ARM and RISC-V: Three Families, Three Histories
- RISC vs CISC: A Real Argument That Stopped Predicting Anything
- Reading Assembly Without Writing It
- Addressing Modes: How an Index Becomes an Address
- Fetch, Decode, Execute — and Why That Story Is Incomplete
- Instruction Fetch: Code Is Data Too
- Decode: Turning Bytes Into Intent
- Execute: Not All Operations Cost the Same
- Load and Store: Why Arithmetic Happens in Registers
- Pipelining: Throughput Without Making Anything Faster
- Pipeline Hazards: The Three Ways Overlap Fails
- Forwarding and Stalls: Paying for Dependencies
- 40/16
Level 4 · Guessing, reordering and doing several things at once
A modern core is a throughput machine: it predicts branches, executes speculatively, runs instructions as their inputs arrive, and only commits in program order. Nearly every mechanism here is a fix for a pipeline hazard from Level 3; SIMD is the other way to do several things at once, one instruction over many elements. This is where the line-by-line mental model dies.
Before moving on: Explain why sorted input runs faster, why IPC rather than clock speed measures a core's work, and say whether a given loop can be vectorized and what stops the compiler when it cannot.
Needs first:Level 3 · The contract, and executing against it- Control Hazards: The CPU Does Not Know Where You Are Going
- Branch Prediction: Guessing Well Enough to Matter
- Misprediction: What a Wrong Guess Costs
- Speculative Execution: Doing Work Before You Know You Need It
- Branchless Code: A Trade, Not an Upgrade
- Out-of-Order Execution
- Dependency Graphs: The Real Shape of Your Code
- Instruction-Level Parallelism
- Superscalar Execution
- Register Renaming
- The Reorder Buffer and Precise State
- IPC: Instructions Per Cycle
- Four Kinds of Parallelism
- SIMD: One Instruction, Many Elements
- Vectorization: Turning a Loop Into Vector Work
- Auto-Vectorization: Verify, Do Not Assume
- 50/14
Level 5 · The memory hierarchy
The deepest level, because this is where most real programs spend their time. Lines, locality, associativity, replacement, thrashing and prefetching — everything behind "why is this loop slow". It starts from the load and store instructions of Level 3 and asks what happens after they leave the core.
Before moving on: Classify a miss as compulsory, capacity or conflict, predict the cliff when a working set outgrows a cache level, and explain why a power-of-two stride can be pathological.
Needs first:Level 3 · The contract, and executing against it- The Memory Hierarchy
- What a Cache Actually Is
- Memory Moves in Lines, Not Variables
- Spatial Locality
- Temporal Locality
- Hits, Misses and What a Miss Actually Costs
- Three Kinds of Miss, Three Different Fixes
- Direct-Mapped Caches: One Address, One Home
- Set-Associative Caches: The Compromise That Won
- Tag, Index and Offset: How an Address Finds Its Line
- Cache Replacement: LRU Is the Idea, Not the Implementation
- Cache Thrashing: Load, Evict, Reload, Repeat
- Working Set: Why Performance Falls Off a Cliff
- Prefetching: The Hardware Guesses What You Will Read Next
- 60/13
Level 6 · Past the cache, and where the bytes sit
DRAM organization, latency against bandwidth as two separate resources, and the layout decisions — alignment, padding, contiguity, AoS versus SoA — that decide how much of each fetched line you actually use. It assumes the cache line from Level 5, because layout only matters in units of lines.
Before moving on: Tell a bandwidth-bound loop from a latency-bound one, compute a struct's padded size from its field order, and explain why an array beats a linked list at equal complexity.
Needs first:Level 5 · The memory hierarchy- Past the Last-Level Cache
- How DRAM Is Organised
- Latency and Bandwidth Are Different Resources
- When the Memory Bus Is the Bottleneck
- When You Cannot Ask the Next Question Yet
- Alignment: Why Addresses Are Not Arbitrary
- Padding: Why Your Struct Is Bigger Than Its Fields
- Endianness: Which Byte Comes First
- What `arr[i]` Actually Compiles To
- Both Are O(n). One Is Far Slower.
- Pointer Chasing: The Address You Do Not Have Yet
- Array of Structs, or Struct of Arrays?
- Data-Oriented Design, Without the Dogma
- 70/9
Level 7 · Address translation and protection
The hardware half of virtual memory: the MMU on the path of every access, the TLB that makes it affordable, and the privilege and permission bits that process isolation is actually built from. It comes after the caches and the layout level because a page-table walk is itself a chain of dependent memory accesses — pointer chasing in the hardware — and the TLB is one more cache with its own misses.
Before moving on: Trace a load through the MMU and TLB, say what a TLB miss costs relative to a cache miss, and explain what changes in the hardware when a system call crosses into the kernel.
- Every Address Your Program Uses Is Fake
- The MMU: Translation and Protection in One Check
- The Page-Table Walk: Dependent Loads All the Way Down
- The TLB: A Cache for Addresses, Not Data
- When Translation Itself Is the Bottleneck
- Huge Pages: More Coverage per Entry, and What It Costs
- What Actually Stops One Process Reading Another's Memory
- Why Kernel Mode Is Actually Privileged
- Why a System Call Costs More Than a Function Call
- 80/18
Level 8 · More than one core
Coherence keeping caches consistent, false sharing as its most expensive surprise, NUMA and affinity — then the ordering rules and atomic instructions that every lock and every lock-free structure is built on. It needs both the caches of Level 5 (coherence is about lines) and the out-of-order core of Level 4 (reordering and store buffers are why a memory model is needed).
Before moving on: Spot false sharing from the layout of a struct two threads update, explain why a store can become visible late on one ISA and not on another, and build a spinlock from compare-and-swap with the right barriers.
Needs first:Level 4 · Guessing, reordering and doing several things at onceLevel 5 · The memory hierarchy- What a Second Core Actually Adds
- Core, Hardware Thread, Software Thread
- SMT: Two Contexts, One Core
- Hardware Threads Are Not OS Threads
- Cache Coherence: Why Shared Memory Works At All
- MESI and Its Relatives
- False Sharing: Independent Data, Shared Line
- NUMA: Not All Memory Is Equally Far
- Thread Affinity: Pinning and Its Price
- Cache Warmth and the Real Cost of Migration
- Sequential Consistency: The Model You Already Have
- Why Your Loads and Stores Happen Out of Order
- Store Buffers: Where Your Writes Wait
- Memory Barriers: Ordering, Not Flushing
- Hardware Memory Models Are Not Language Memory Models
- Atomic Instructions: What the Hardware Actually Guarantees
- Compare-and-Swap: The Primitive Everything Is Built On
- What a Mutex Actually Does
- 90/36
Level 9 · The rest of the machine, and reading it
Interrupts, DMA and the path to devices; GPUs and accelerators as throughput hardware with a transfer bill; the counters that tell you what the machine is doing; and where all of it shows up in code you already write. It comes last because reading a counter, a side channel or a virtual machine only makes sense once you know the speculating core, the translation hardware and the coherent caches it is reporting on.
Before moving on: Read performance counters to say whether a program is compute-bound or memory-bound, decide whether a kernel is worth moving to a GPU once transfers are counted, and explain Spectre as a consequence of speculation.
Needs first:Level 4 · Guessing, reordering and doing several things at onceLevel 7 · Address translation and protectionLevel 8 · More than one core- Interrupts: How Hardware Gets the CPU's Attention
- Polling versus Interrupts
- DMA: Moving Bytes Without the CPU
- I/O Architecture: The Interconnect Is a Shared Resource
- Memory-Mapped I/O: When a Store Is Not a Store
- PCIe: Lanes, Generations and the Transfer Budget
- The Storage Path: Why One Small Read Is the Worst Case
- CPU Cache Is Not the Page Cache
- CPU or GPU: Two Bets About What Work Looks Like
- What Is Actually Inside a GPU
- Lanes, Divergence and Coalescing
- The Transfer You Forgot to Count
- What GPU-Friendly Work Has in Common
- Accelerators: The Specialization Spectrum
- The Specialization Trade-off
- LLM Inference Is a Memory Bandwidth Problem
- Model Memory, and Why the Naive Number Is Always Too Low
- The CPU Counts Itself
- CPI and IPC: The Number Everyone Misreads
- Busy Is Not the Same as Working
- Misses That Overlap Are Nearly Free
- Your Code Is Data Too
- Every Way a CPU Microbenchmark Lies
- The Clock Is a Variable
- The First Ten Seconds Lie
- Performance Per Watt
- Throughput Improved, Latency Did Not
- Cache-Aware Algorithms
- Matrix Tiling: Same Arithmetic, Ten Times Faster
- The Compiler Reordered It Before the CPU Did
- Why Reading the Source Cannot Tell You the Cost
- Side Channels: When Performance Optimisations Leak
- Spectre and Meltdown: When Speculation Crossed a Boundary
- The Hardware That Makes Virtual Machines Possible
- What a vCPU Actually Is
- From malloc to Cache Lines