intermediate
Two Million and Climbing
Read the evidence before you read the options. The signals are presented the way a dashboard would present them — nothing is labelled with the answer.
The report
Order confirmation emails are going out hours late. The queue depth chart has been climbing since about 06:00 and is now at 2.1 million. The team is in a thread arguing about how many workers to add, and someone has already scaled the worker pool from 40 to 120 with no effect.
The system
Queue and workers — currentILLUSTRATIVE
| Signal | Value | What it tells you |
|---|---|---|
| Arrival rate | 9,800 msg/s (last week: 9,600) | Messages arrive at essentially the same rate as last week. |
| Completion rate | 6,150 msg/s (last week: 9,700) | Workers are finishing about a third fewer messages than they were. |
| Queue depth | 2.1 M, growing ~3,650/s | The backlog grows at the difference between arrival and completion. |
| Oldest message age | 4 h 12 min | The message at the head of the queue was enqueued four hours ago. |
| Worker count | 120 (was 40 at 07:30) | The pool was tripled two hours ago. |
| Worker CPU | 11% | Worker processes are almost entirely idle. |
| Completion rate before scaling | 6,050 msg/s at 40 workers | Tripling the worker count changed throughput by under 2%. |
Email provider — as seen from our workersILLUSTRATIVE
| Signal | Value | What it tells you |
|---|---|---|
| Send call p50 | 840 ms (was 45 ms) | The typical provider call takes nearly a second, up from under fifty milliseconds. |
| Send call p99 | 2.9 s (was 180 ms) | The tail of provider calls has grown by more than an order of magnitude. |
| Provider error rate | 0.3% | Calls are succeeding; they are simply taking longer. |
| Concurrent in-flight sends | 118 | Nearly every worker is inside a provider call at any moment. |
What is the constraint?