Data Engineering Roadmap

Ten levels, each defined by what you can actually build once you have it rather than by what you have read. The order matters: every level assumes the failure modes of the one before it, and the last one is where a number becomes explainable.

0 / 244 mastered0%
Level 1

The shape of the problem

You can explain why analytics moved off the production database, and answer an analytical question without taking anything down.

What Data Engineering Actually Is
The Fundamental Data Journey
OLTP Workloads
OLAP Workloads
OLTP vs OLAP
Workload Isolation
The OLTP to OLAP Journey
ETL: Transform Before the Data Lands
ELT: Load First, Transform Where the Data Lives
ETL vs ELT: Choosing by Constraint, Not by Fashion
Who Actually Consumes This Data
Source of Truth
Level 2

Where data lives

You can choose between a lake, a warehouse and a lakehouse for a real workload, and say what each one costs you.

Row vs Column Storage
Columnar Execution
Why Analytical Data Compresses
Dictionary, Run-Length, Delta and Bit Packing
Parquet
Parquet Internals
The Parquet Read Path
Avro
ORC
Parquet vs Avro
CSV, JSON and Their Limits
The Data Lake
The Data Warehouse
The Lakehouse
Open Table Formats
Lake vs Warehouse vs Lakehouse
Object Storage as Data Infrastructure
Separating Storage from Compute
Data Marts
Level 3

Getting data in

You can build an ingestion path that survives a failure without losing or duplicating a day, and explain the predicate it extracts on.

Data Ingestion
Ingestion Sources
Batch Ingestion
Incremental Extraction
Streaming Ingestion
Batch vs Streaming Ingestion
Ingestion Failure & Recovery
The Raw Landing Zone
Raw, Staging, Curated: Layers by Purpose
Medallion: One Naming Convention Among Several
Where the Transformation Actually Runs
Level 4

Modelling for questions

You can design a model whose grain is declared, whose history is preserved where it needs to be, and whose measures do not multiply when joined.

Analytical Data Modeling
Operational vs Analytical Models
Fact Tables
Grain: What Does One Row Represent?
Dimension Tables
Surrogate Keys
Star Schema
Snowflake Schema
Slowly Changing Dimensions
SCD Type 2 in Practice
Snapshot Tables
Event vs Snapshot Modeling
Level 5

Transformation and orchestration

You can express a platform as a tested dependency graph, and re-run any part of it without corrupting what is already correct.

Data Transformation
SQL Transformations
dbt Concepts
The Transformation DAG
DAGs in Data Pipelines
Topological Execution
Model Layering
The Metrics Layer
Orchestration
Scheduler vs Orchestrator
Airflow Concepts
Task Dependencies
When a Task Fails Mid-DAG
Idempotent Data Pipelines
Incremental Processing
The High-Water Mark
Level 6

Physical layout

You can make a query read a fraction of what it used to, and explain exactly why — the highest-leverage skill in analytics.

Physical Data Layout
File Size and the Small-Files Problem
File Compaction
Partitioning
Partition Pruning
Partition Cardinality
Clustering and Sort Order
Bucketing
The Partitioning Decision
Query Engines
Distributed Query Execution
Predicate Pushdown
Projection Pushdown
Source Pushdown
Vectorized Execution
Level 7

Change and streams

You can capture changes from a database log, publish them to a replayable stream, and reason correctly about event time and lateness.

Change Data Capture
CDC vs Polling
What a CDC Event Contains
CDC Ordering and Transaction Boundaries
Snapshot and Stream: the Bootstrap Problem
CDC Failure Modes and the Retention Deadline
CDC and Schema Drift
The Event Log
Message Brokers: Log-Shaped and Queue-Shaped
Kafka as a Log, Not a Queue
Topics and Partitions
Event Keys and Partition Assignment
Consumer Groups and the Parallelism Ceiling
Retention and Replay
Offsets and Commits
Stream Processing
Stateless Stream Processing
Stateful Stream Processing
Streaming State
Event Time
Processing Time
Ingestion Time
Late Events
Windows
Tumbling Windows
Sliding Windows
Session Windows
Watermarks
Stream Joins
Level 8

Distributed compute

You can read an execution plan, find the shuffle, spot the skew, and know why adding workers will not help.

Distributed Data Processing
The Spark Execution Model
Partitions: the Unit of Parallelism
Stages and Tasks
The Shuffle
Narrow and Wide Transformations
Data Skew
Straggler Tasks
Salting a Skewed Key
Broadcast Joins
Lazy Evaluation
Query Optimizers
Flink Concepts
Federated Query
Batch and Streaming Unification
Level 9

Trust

You can say whether a number is trustworthy and prove it — and find out that it is not before a finance team does.

Data Quality
The Dimensions of Data Quality
Data Tests
Distribution Tests
Freshness Checks
Reconciliation
Quality Alerting
The Data Quality Dashboard
Who Owns Data Quality
Data Contracts
Schema Evolution
Backward Compatibility
Forward Compatibility
Schema Registry
Breaking Schema Changes
Semantic Changes
Nullability & Defaults
Contract Enforcement
Metadata: Technical, Operational and Business
The Data Catalog
Data Lineage
Column-Level Lineage
Impact Analysis
Data Ownership
Data Discovery
Dataset Documentation
Data Observability
Pipeline Observability
Pipeline Metrics
Freshness Monitoring
Volume Anomalies
Data Incidents
Debugging a Data Incident
Lineage Debugging
Level 10

Production data engineering

You can run a platform: recover it, govern it, pay for it, and explain any number on it back to its source.

Backfills
What Backfills Break
Planning a Backfill
Validating a Backfill Before You Publish
Reprocessing vs Retrying
Late-Arriving Data
Deduplication
Upserts and Merges
Full Refresh vs Incremental
Replay from the Log
Pipeline Reliability
Atomic Publish
Checkpointing
Partial Failure
Retries in Pipelines
Pipeline SLOs
The Freshness SLO
Rolling Back Data
Data Governance
Data Classification
PII in Pipelines
Data Minimization
Data Retention
Data Access Control
Row and Column Security
Data Masking, Tokenisation & Encryption
Deletion Requests
What Actually Drives Data Platform Cost
Scan Cost
Compute Waste
Storage Lifecycle
Cost Attribution
Cost vs Freshness
Data Architecture Patterns
The Central Warehouse
The Event-Driven Data Platform
Lambda Architecture
Kappa Architecture
Data Mesh
Data Products
Data Platform Engineering
The Self-Service Data Platform
Cloud Data Services
Comparing Analytical Warehouses
BigQuery Concepts
Snowflake Concepts
ClickHouse Concepts
DuckDB Concepts
Managed Streaming Platforms
Choosing an Analytical Platform
Data Engineering for Agents
The LLM Data Pipeline
Chunking Pipelines
Embedding Pipelines
Re-embedding
Vector Data Engineering
Evaluation Data Pipelines
Agent Observability Data
Feature Pipelines
Where Did This Number Come From?
Two Dashboards, Two Numbers
Missing Rows
Duplicate Rows
Stale Dashboards
The Pipeline Succeeded. The Data Is Wrong.
Data Engineering Anti-Patterns
Data Platform Anti-Patterns
Data Engineering and Database Engineering
Data Engineering and Distributed Systems
Data Engineering and Backend Engineering
Data Engineering and Cloud Infrastructure
Data Engineering and DevOps
Data Engineering and Observability
Data Engineering and Security
Data Engineering vs Its Neighbours
The Data Loop
Trusting Data
What Goes Wrong Between Source and Dashboard
Data Engineering and Database Engineering