Operating Systems Roadmap
Start at Programs and follow one program down to the CPU, RAM, disk and NIC. Every stage names what it needs first and what you should be able to do before moving on; the last stage adds the OS + Networking lessons. Progress is stored locally in your browser.
Where to start
Operating Systems
15 stages · 0/64 lessonsWhat actually happens between your program and the hardware: processes, threads, the scheduler, syscalls, memory, files, I/O, concurrency and containers.
- 10/4
Programs
Start hereWhat an OS is for, and what happens between
./serverand the first instruction on a CPU: the loader, the address space, and why one executable on disk is not the same thing as the three instances running from it. Everything else in the domain is a layer under this one picture.Before moving on: Trace
./serverfrom a file on disk to a running process with its own address space, and explain why the executable and the process are different things. - 20/3
Processes
A process has an identity, an address space, a state, open resources and a parent. The process table, the state machine from Runnable to Zombie, and
fork+execas the two halves of "start a program". Threads, scheduling and IPC all assume this unit.Before moving on: Draw the process state machine, explain what
forkandexeceach do, and say what a zombie is and who reaps it.Needs first:Programs - 30/7
Threads
Several flows of control inside one address space; concurrency as interleaving versus parallelism as simultaneous cores; and how C++, JavaScript/TypeScript and Python each map "do many things at once" onto OS threads, event loops and async I/O.
Before moving on: Choose between threads, async I/O and processes for a given workload, and explain why an event loop on one core is concurrency but not parallelism.
Needs first:Processes - 40/3
Scheduling
A hundred runnable tasks and eight cores: ready queues, time slices, priority and preemption, and what a context switch actually saves, restores and throws away (registers, the TLB, the cache). This is the mechanism that makes the interleavings from the Threads stage real.
Before moving on: Predict which task runs next under a given policy, and list what a context switch saves, restores and invalidates.
- 50/2
System Calls
Why a program cannot write to the disk itself: the user/kernel boundary, the trap into kernel mode, what a syscall costs, and why "mostly system time" in
topis a diagnosis in itself. Files, I/O and sockets are all reached through this door.Before moving on: Explain why a program cannot write to the disk itself, walk through what happens on a trap into kernel mode, and read a high system-time number in
topas a diagnosis. - 60/4
Files / Descriptors
A filename is not a file: paths, directories, metadata and permissions; the per-process descriptor table (and what FD 3 is); how a path becomes blocks on storage; and inodes on Unix-style systems. Descriptors are the handle I/O, pipes and sockets all share.
Before moving on: Explain the difference between a filename, an inode and a descriptor, say what FD 3 is in a fresh process, and follow a path down to blocks on disk.
- 70/6
Virtual Memory
Where a local variable lives, where an object lives, why recursion has a limit and what
malloc/new/an object allocation actually asks the kernel for — then the illusion underneath: every process believes it owns a huge private, contiguous memory.Before moving on: Say where a local variable, an object and a stack frame live, explain what limits recursion depth and what
mallocasks the kernel for, and state what every process is promised about its own address space. - 80/7
Paging
How the illusion is built: pages and frames, multi-level page tables, the TLB that makes translation affordable, page faults as the normal mechanism (not an error), memory pressure and swapping,
mmap, and copy-on-write behind a fastfork().Before moving on: Translate a virtual address through a page table by hand, explain when a page fault is normal and when it is a problem, and say why
fork()is cheap thanks to copy-on-write.Needs first:Virtual Memory - 90/4
I/O
Follow one
read()from the call through the page cache to the SSD and back; blocking, non-blocking, asynchronous and multiplexed I/O;select/pollversusepoll/kqueue; and why a file, a pipe and a socket are the same kind of thing to the kernel. The event loop from the Threads stage is built on the multiplexing here.Before moving on: Follow a
read()from the call through the page cache to the SSD and back, and choose between blocking, non-blocking, async and multiplexed I/O for a server holding 10,000 sockets. - 100/3
Concurrency
Two threads add one to a counter and the result is one: the zoo of concurrency bugs, why a race is a property of the interleaving and not of the code, and the critical section as the thing every fix protects. It comes after Scheduling because the scheduler is what chooses the interleaving.
Before moving on: Find the interleaving that makes two increments produce one, mark the critical section in a piece of code, and explain why the bug belongs to the schedule rather than to the code.
- 110/4
Synchronization
The tools that make a critical section safe — mutexes, counting semaphores, atomic compare-and-swap — and the four ingredients that turn a lock into a deadlock, with the wait-for graph that finds it.
Before moving on: Pick a mutex, a semaphore or an atomic for a given critical section, and find a deadlock in a wait-for graph and name which of the four conditions to break.
Needs first:Concurrency - 120/4
IPC
Two isolated processes need to talk: pipes, shared memory, message queues, signals and sockets compared on speed, isolation, complexity and local-versus-remote. Pipes are descriptors and shared memory is mapped pages, so this stage reuses Files and Paging.
Before moving on: Pick pipes, shared memory, a message queue, signals or a socket for a given pair of processes, and say what each costs in speed, isolation and complexity.
- 130/1
Sockets
What the kernel actually hands you when you call
socket(): a descriptor with send and receive buffers behind it, a state machine, and the bridge into the Networking domain. It is the last IPC option, seen from the I/O side.Before moving on: Describe what
socket()returns, name the buffers and the state behind it, and explain why a socket is read and written like a file. - 140/3
Containers
A container has no kernel of its own, so what isolates it? Namespaces, control groups and layered filesystems on a shared host kernel — and how that differs from a hypervisor running a whole guest OS. Every mechanism here is one of the earlier stages with a boundary drawn around it.
Before moving on: Explain what isolates a container without a kernel of its own — namespaces, cgroups and a layered filesystem — and say when a VM is the right choice instead.
- 150/9
Performance / Debugging
Break a simulated OS on purpose, then learn the playbook for high CPU, high memory, hangs and
EMFILE; the capstone explains a 50,000-connection server layer by layer, and thetogetherlessons followsend()across the wire to arecv()on another machine. It is last because every diagnosis names one of the earlier mechanisms.Before moving on: Diagnose high CPU, high memory, a hang and
EMFILEfromtop,ssandstraceoutput, and explain a 50,000-connection server layer by layer fromsend()torecv().- The OS Simulator: Cores, Processes, RAM, I/O and Locks
- Break the OS: Predict, Break, Diagnose
- The OS Debugging Playbook: Four Symptoms, Fourteen Causes
- Capstone: A Server With 50,000 Concurrent Connections
- Follow send() Through the OS to recv()
- Build a Tiny Server: V0 to V5
- C10K: Ten Thousand Connections, Then a Million
- Combined Failure Simulator: Break a Layer, Watch It Propagate
- Capstone: Three Seconds from Warsaw