Comparisons
Two things routinely conflated, put side by side. Neither column is the winner — what decides is the workload and the machine.
CPU vs GPU
Latency-optimized versus throughput-optimized. A CPU spends transistors on making one thread fast — speculation, big caches, deep out-of-order machinery. A GPU spends them on running thousands of lanes of regular work at once, and hides latency with parallelism instead of prediction.
Fast on serial, branchy, irregular and latency-sensitive work
Limited parallelism; expensive per unit of throughput
General computation, control-heavy logic, low-latency response
Enormous throughput on wide, regular, arithmetic-heavy work
Transfer cost, divergence penalties, and uselessness on serial work
Dense linear algebra, graphics, ML, image and signal processing
| Dimension | CPU | GPU |
|---|---|---|
| Hides latency by | Speculation, caches, out-of-order execution | Massive thread parallelism |
| Branchy code | Handled well by the predictor | Divergent lanes serialise |
| Data must move | No | Usually yes, and it often dominates |
| Good at small work | Yes | No — launch and transfer overheads swamp it |