Backend Engineering Roadmap

Nine levels, each defined by what you can build once you have it rather than by what you have read. The order matters: every level assumes the failure modes of the one before it.

0 / 200 mastered0%
Level 1

Request in, response out

You can build an HTTP service that accepts a request, rejects malformed input at the edge, runs a handler and returns a correct status code — and you can explain what happened between the socket and the response body without pointing at a framework.

What a Backend Actually Is
The Request Lifecycle
What the Backend Is Responsible For
The Trust Boundary
Anatomy of an HTTP Server
Request and Response Objects
Status Codes From the Server's Side
Request Bodies and Streaming
How a Route Becomes a Function Call
Path Parameters
Query Parameters
Route Precedence
What a Handler Is Responsible For
The Middleware Pipeline
The Three Validations
Transport Validation
Reporting Validation Failures
Level 2

Structure, data access and who is calling

You can build a service whose handlers are thin, whose business rules live somewhere testable, and which knows both who the caller is and what that caller is allowed to touch. This is the level at which a personal project becomes something other people can safely log into.

Transport, Application, Domain, Infrastructure
The Service Layer
The Repository Layer
Fat Controllers
Dependency Management Without the Container
Business Validation
Parse, Do Not Validate
Choosing a Data Access Layer
What an ORM Actually Does
Middleware Ordering Is a Correctness Decision
Authentication in a Backend
Credentials and Password Handling
Session Authentication
Token Authentication and the Revocation Problem
Authorization in Backends
Authentication vs Authorization
Where the Check Belongs
Role-Based Access Control
Object-Level Authorization
Level 3

Transactions, queries, caches and other people's APIs

You can decide what belongs in one atomic unit, write queries that do not multiply with the result set, add a cache that does not lie, and call a third party without your own service hanging when theirs does.

Transactions from Application Code
Where the Transaction Boundary Goes
One Transaction or Two
External Calls Inside a Transaction
What an ORM Buys and What It Costs
The N+1 Query Problem
Eager Loading and Batching
Query Builders
Raw SQL in Application Code
Connection Pools
Caching in Backends
Cache-Aside
Cache Invalidation
TTL and Expiry
When Not to Cache
Calling Something You Do Not Control
Timeouts
Serialization: Objects to Bytes
Three Models, Not One
Schema Leakage
Level 4

Work that outlives the request

You can move work off the request path deliberately — jobs, queues, events, inbound webhooks and uploads that never touch your process — and answer "what does the user see while that is still running?".

Request or Background?
Background Jobs
Job Queues
Queue Semantics
Scheduled Jobs
Dead-Letter Queues
Commands vs Events
Event-Driven Backends
Naming Events
Writing Event Consumers
Inbound Webhooks
Webhook Signature Verification
Outbound Webhooks
Choosing an Upload Path
File Uploads Through the Backend
Presigned URLs
Object Storage
What Happens After the Bytes Land
Serving Files
Level 5

Timeouts, retries, idempotency and limits

You can build a service that survives a flaky dependency and a client that presses the button twice — because every retryable path is also a safe-to-retry path, and every unbounded thing has a bound.

Retries
Backoff and Jitter
Circuit Breakers
Bulkheads
Rate Limiting
Rate Limit Algorithms
Idempotency in Backends
Idempotency Keys
The Idempotency Key Flow
Idempotency Storage
At-Least-Once Delivery
Duplicate Detection
Job Idempotency
Webhook Idempotency
Webhook Retries and Ordering
Backend Races
Optimistic Concurrency
Pessimistic Locking
Atomic Operations
Resource Limits
Level 6

Observability, testing, security and performance work

You can answer questions about production that you did not anticipate when you deployed, prove a change is safe before shipping it, and pass a security review of the service rather than of the network in front of it.

An Error Taxonomy That Maps Cause to Response
Error Boundaries: Three Translations, Not One
Not Leaking Your Internals
Correlation Ids That Survive Every Hop
What a Backend Should Actually Log
Structured Logging
The Metrics a Backend Must Emit
Tracing From the Backend's Side
Health Checks: Startup, Readiness, Liveness
Configuration: Separating Code From Environment
Secrets Are Not Configuration
Validate at Startup, Fail Loudly
A Test Strategy Chosen by What Each Layer Can Prove
Test Against the Real Database
Contract Tests Between Services
Performance Testing a Backend
The Backend Security Checklist
SQL Injection
Command Injection
SSRF — When the Backend Fetches a URL
Secrets in Logs
What Serialization Costs
Level 7

Scaling out and shipping without dropping requests

You can run more than one copy of your service behind a load balancer, deploy a new version while the old one is still serving, and migrate a schema that two versions are reading at the same time.

Stateless Services
Making an Existing Service Stateless
Horizontal vs Vertical Scaling
Load Balancing, From the Backend's Side
Sticky Sessions
Autoscaling a Backend
Read Replicas From the Application
Pagination That Survives a Large Table
Deployment Models
Containerizing a Backend
Graceful Shutdown
Rolling Deployments
Expand and Contract Migrations
Schema Migrations from the Application Side
Running Two API Versions in One Service
Mapping Services Across Cloud Providers
Running a Backend on Kubernetes
Serverless Backends
Choosing a Runtime
Backend Runtime Models
Level 8

Distributed patterns, failure handling and multi-tenancy

You can reason about a system where a function call has become a network call: what happens when half of a write succeeds, how one slow dependency takes down three services that do not depend on it, and how one deployment serves customers who must never see each other.

The Monolith
The Modular Monolith
Microservices
Comparing Backend Architectures
Synchronous vs Asynchronous Communication
Failure Propagation
Cascading Failure
The Dual Write Problem
The Transactional Outbox
Eventual Consistency in Practice
Keeping a Search Index in Sync
Multi-Tenancy
Tenant Isolation
Attribute-Based Access Control
Backpressure
Worker Scaling
Local vs Distributed Cache
Cache Stampede
Unbounded Concurrency
Request Coalescing
Deadlocks in Application Code
Level 9

Production backend engineering

You can be the person who is called when the API was 100 ms yesterday and is 3 s today: reason from symptom to cause, ship the mitigation safely, and hold the operational judgement about what to change and what to leave alone.

Debugging a Backend in Production
Why Is My API Slow?
The Common Backend Failures
Backend Code Smells
Deploys Are the First Suspect
Memory Leaks in Backend Services
Retry Storms
Connection Pool Exhaustion
Queue Backlog
Blocking the Event Loop
The Node Event Loop
Python Runtime Models
Worker Processes
C++ Backend Services
Canary Deployments
Blue-Green Deployments
Feature Flags: Rollout, Kill Switches and Debt
Defence in Depth
Dependency Security
Keep-Alive and Connection Reuse
Accepting Connections
Parsing HTTP
Deserialization: Bytes to Objects
Every Input Surface
Database Constraints
When the Repository Is Just Indirection
Alternatives to Layering
What Belongs in the Pipeline
Request Context Propagation
The Error Boundary
Authenticate First, or Rate-Limit First?
Where Sessions Live
OAuth and OIDC From the Backend Side
API Keys
Email and Notifications
Backend and Its Neighbours
The Backend Reasoning Loop
What an Agent Adds to a Backend
A Tool Call Is a Backend Call
Agent Authorization
Budgets, Deadlines and Step Limits
Agent Audit Logs