Performance Roadmap
One track, nine levels. Start at Level 1 with measuring before you touch anything; every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.
Where to start
Observability & Performance
9 stages · 0/119 lessonsReading what a production system is doing, finding the bottleneck, and proving that a change actually helped.
- Level 1 · Measure before you touch anything
- Level 2 · Read the numbers honestly
- Level 3 · Find where the time goes
- Level 4 · Understand why loaded systems get slow
- Level 5 · Identify the constrained resource
- Level 6 · Diagnose the data layer
- Level 7 · Know your runtime and your users
- Level 8 · Produce numbers that mean something
- Level 9 · Run it in production, at scale
- 10/9
Level 1 · Measure before you touch anything
Start hereThe discipline that separates diagnosis from guessing: a symptom, a signal that confirms it, and a hypothesis written down before you look. Learn what each signal type can and cannot tell you, the four numbers that describe any service, and how instrumentation reaches a backend at all. Everything after this assumes you measure first.
Before moving on: Turn a vague complaint into a symptom, the signal that would confirm it and a written hypothesis, and say whether a metric, a log or a trace answers the question in front of you.
- 20/13
Level 2 · Read the numbers honestly
Averages hide outages and a user id in a label can take down the metrics backend. Counters, gauges and histograms as different questions, percentiles as the only honest summary of latency, and logs structured well enough to reconstruct a failure. It comes second because every later stage reads a
p99or a log line and needs you to read it right.Before moving on: Read a p50 / p99 pair and say which one the user feels, pick counter, gauge or histogram for a new measurement, and reject a label that would explode cardinality.
Needs first:Level 1 · Measure before you touch anything- Four Metric Types, Four Questions
- Counters: The Slope Is the Signal
- Gauges: Blind Between Scrapes
- Histograms: A Distribution You Can Afford to Keep Forever
- The Average Was Fine and Users Were Not
- Percentiles: Which One, and How Many Users Is That?
- Cardinality: The Label That Took Down Monitoring
- Label Sets That Survive a Year
- Structured Logging: Fields a Program Can Read
- Log Levels Are a Convention, Not a Standard
- Correlation IDs: Turning Lines Into a Story
- What You Just Wrote Into a Log Half the Company Can Read
- The Log Bill and What It Is Buying
- 30/14
Level 3 · Find where the time goes
Tracing locates time across services; profiling locates cost inside a process. Read a waterfall, find the critical path, spot N+1 by its shape, and know which of the two tools answers the question in front of you. These are the two tools every diagnosis from here on reaches for.
Before moving on: Read a trace waterfall, mark the critical path, spot an N+1 by its shape, and decide from a flame graph whether a process is CPU-bound or waiting on I/O.
- Where the Request Actually Went
- Trace, Span, Attribute, Status
- Parents, Children and Links
- Carrying the Trace Across the Gap
- Reading the Waterfall
- The Critical Path Is the Only Path That Pays
- The Comb: N+1 as a Visible Shape
- Sampling Without Throwing Away the Evidence
- When the Trace Runs Out of Answers
- Self Time, Total Time, and Where the CPU Went
- Reading a Flame Graph
- Allocation Rate Is a Cost Even Without a Leak
- Computing or Waiting?
- Always-On Profiling, and the Diff That Finds Regressions
- 40/9
Level 4 · Understand why loaded systems get slow
The physics. Queueing makes systems slow long before they fail, tail latency dominates user experience at scale, and Little's Law sizes every pool you will ever configure. It sits after percentiles because the queueing curve only makes sense once you read latency as a distribution, and before the resource and data-layer stages because saturation is what those diagnose.
Before moving on: Size a pool with Little's Law, explain why latency climbs long before utilisation reaches 100%, and set a timeout from a latency budget instead of a guess.
Needs first:Level 2 · Read the numbers honestly- Latency Is a Distribution, Not a Number
- Tail Latency: Why p50 Being Fine Does Not Help
- Latency Budgets: Spending 200 Milliseconds on Purpose
- Throughput: The Number That Means Nothing Without a Latency Bound
- Little's Law as Working Intuition
- Queueing: Why Systems Get Slow Before They Get Broken
- Saturation: The Reading Utilization Cannot Give You
- Concurrency Limits: An Unbounded Server Is a Slower Server
- Timeouts: The Latency Contract Nobody Writes Down
- 50/8
Level 5 · Identify the constrained resource
CPU, memory, disk and network, each with the signal that identifies it as the constraint — and the difference between a memory leak and a cache nobody bounded. Profiling tells you where cost goes; this stage tells you which resource has run out, so it follows profiling and saturation.
Before moving on: Name the constrained resource from its signal alone, and tell a memory leak from an unbounded cache by plotting memory against traffic over a full daily cycle.
- What "CPU Is At 60%" Actually Means
- CPU Saturation: When Cores Become the Queue
- Algorithmic Cost in a Request Handler
- Reading Memory: RSS, Heap, Working Set and the Number on Your Dashboard
- Memory Leaks: Growth That Does Not Come Back
- Leak or Unbounded Cache? The Question That Picks the Fix
- Disk and Storage: Latency, Throughput, IOPS and the fsync Tax
- Network Signals: Is It the Network, or the Service on the Other End?
- 60/15
Level 6 · Diagnose the data layer
Databases, caches and queues are where most production latency actually lives. Slow-query workflow, lock waits with idle CPU, pool saturation, hit rates that lie, stampedes, hot keys, backlogs and retry storms. Every one of these is queueing or a constrained resource seen from outside the process, which is why it comes after both.
Before moving on: Run the slow-query workflow to a fix, explain lock waits on an idle CPU, size a connection pool, and read queue depth and oldest-message age as two different alarms.
Needs first:Level 4 · Understand why loaded systems get slowLevel 5 · Identify the constrained resource- Which Signal Actually Means "The Database Is Slow"
- The Slow Query Workflow
- An Index Scan Is Not Automatically Faster
- When the Join Strategy Is the Bottleneck
- Low CPU, High Latency: Lock Contention
- Connection Pool Saturation: Waiting in Front of an Idle Database
- Replication Lag: Reads That Are Correct and Stale
- A 95% Hit Rate Tells You Almost Nothing
- Cache Stampede: Everyone Misses at Once
- Hot Keys: When Aggregate Metrics Hide a Saturated Node
- Six Queue Signals, Two That Wake You Up
- The Backlog Arithmetic: Four Levers and a Drain Time
- Depth Is Not an Emergency; Age Is
- Retry Storms: The Load You Generated Yourself
- Twenty Workers, All Busy, Five Hundred Waiting
- 70/12
Level 7 · Know your runtime and your users
Garbage collection, event loops, interpreters and JIT warm-up — labelled per runtime because none of it generalizes. Then the user's half of the budget: the browser waterfall, bundle cost and the vitals that measure experience. Both halves read a profile or a waterfall, so profiling and the resource stage come first.
Before moving on: Attribute a tail spike to garbage collection or event-loop lag from the right runtime signal, and read a browser waterfall to say whether the page is slow or the API is.
- Garbage Collection: Pause, Throughput, Footprint — Pick Two
- Event-Loop Lag: One Callback, Everybody Waits
- JavaScript Runtime Performance: V8 Where It Costs
- CPython Performance: The Interpreter Tax and the GIL
- C++ Memory Performance: Allocation, Copies and Locality
- JIT and Warm-Up: The First Thousand Requests Are a Different Program
- The Half of the Budget You Cannot See From the Server
- Core Web Vitals as Signals, Not Scores
- Reading the Browser Waterfall
- JavaScript Costs Four Times, Not Once
- Images: The Largest Bytes, Rarely the Largest Block
- Layout, Paint and the Main Thread
- 80/13
Level 8 · Produce numbers that mean something
Load-test shapes that answer different questions, coordinated omission, benchmark hygiene, and telling a real regression from noise. Then capacity: headroom, autoscaling that arrives in time, and what a request actually costs. A load test is queueing theory run on purpose and a capacity estimate is arithmetic on percentiles, so this waits for both.
Before moving on: Pick the load-test shape that answers the question you have, run a benchmark that avoids the common fallacies, and state a capacity estimate with its headroom and its assumptions labelled.
- Load Testing: What Question Is This Test Answering?
- Load Test Shapes: The Shape Is the Hypothesis
- Coordinated Omission: When the Load Generator Lies
- Benchmarking: Does This Number Answer My Question?
- Benchmark Fallacies: Confident Numbers That Are Wrong
- Microbenchmark or End-to-End: Why p99 Did Not Move
- Regression or Tuesday? Telling a Real Change from Noise
- Capacity Planning: Traffic to Machines
- Headroom: The Capacity You Deliberately Do Not Use
- Autoscaling: Scaling on the Right Signal
- Autoscaling Lag: The Gap Where the Outage Lives
- Cost per Request: The Other Performance Metric
- Capacity or Efficiency: Which Problem Are You Solving?
- 90/26
Level 9 · Run it in production, at scale
SLOs that describe user experience, alerts worth waking up for, error budgets as a decision tool — then distributed tail amplification, cross-region physics, incident method, agent latency budgets, and the trade-offs that make a system faster but worse. It comes last because an SLI is a percentile with a target, an incident is every earlier stage under pressure, and a trade-off is only visible once you can measure both sides.
Before moving on: Write an SLI and SLO for one user journey, set a burn-rate alert, and run an incident from the pivotal signal to a review that names the regression test and the new alert.
- SLIs: Measuring What the User Actually Feels
- SLOs: A Target, a Window, and a Reason
- SLAs: The Promise With Money Attached
- Error Budgets: Unreliability You Are Allowed to Spend
- Alerts Worth Waking Someone For
- Alert Fatigue: The Page Nobody Reads
- Burn-Rate Alerts: How Fast Is the Budget Going?
- Dashboards Built Around Questions
- What Changes When Work Crosses a Machine
- Fan-Out: Waiting for the Slowest of Seven
- Sequential or Parallel: Same Work, Different Latency
- Cross-Region Latency Is Physics, Not Configuration
- Packet Loss Buys You a Timeout, Not a Retransmit
- Agreement Costs Round Trips
- Debugging an Incident in Progress
- Correlation Is Not the Root Cause
- Reading a Timeline: Observation Order Is Not Causal Order
- "What Changed?" — Deploy Markers and the Invisible Deploys
- The Bottleneck Moves After Every Fix
- Every Optimization Buys Something and Sells Something
- Performance and Observability Anti-Patterns
- Where an Agent Run Actually Spends Its Time
- Inside One Model Call: Queue, First Token, Generation
- Reading an Agent Run as a Trace
- What One Agent Run Costs, and Which Term Dominates
- The Agent Returned 200 OK and the Answer Was Wrong