intermediate

Two Million and Climbing

Read the evidence before you read the options. The signals are presented the way a dashboard would present them — nothing is labelled with the answer.

The report

Order confirmation emails are going out hours late. The queue depth chart has been climbing since about 06:00 and is now at 2.1 million. The team is in a thread arguing about how many workers to add, and someone has already scaled the worker pool from 40 to 120 with no effect.

The system
9,800 msg/sOrder APIorders.confirmWorkers ×120Email provider
Queue and workers — currentILLUSTRATIVE
SignalValueWhat it tells you
Arrival rate9,800 msg/s (last week: 9,600)Messages arrive at essentially the same rate as last week.
Completion rate6,150 msg/s (last week: 9,700)Workers are finishing about a third fewer messages than they were.
Queue depth2.1 M, growing ~3,650/sThe backlog grows at the difference between arrival and completion.
Oldest message age4 h 12 minThe message at the head of the queue was enqueued four hours ago.
Worker count120 (was 40 at 07:30)The pool was tripled two hours ago.
Worker CPU11%Worker processes are almost entirely idle.
Completion rate before scaling6,050 msg/s at 40 workersTripling the worker count changed throughput by under 2%.
Email provider — as seen from our workersILLUSTRATIVE
SignalValueWhat it tells you
Send call p50840 ms (was 45 ms)The typical provider call takes nearly a second, up from under fifty milliseconds.
Send call p992.9 s (was 180 ms)The tail of provider calls has grown by more than an order of magnitude.
Provider error rate0.3%Calls are succeeding; they are simply taking longer.
Concurrent in-flight sends118Nearly every worker is inside a provider call at any moment.
What is the constraint?