The Deploy That Removed Its Own Capacity
A routine deploy started at 16:05. By 16:15 the service is serving 502s for about a third of requests, the fleet is down from 12 instances to 4, and the deployment is still "in progress". A rollback was requested at 16:16 and it is also stuck. Latency on the four surviving instances is eight seconds.
The infrastructure under investigation
Exposure is shown, because half of these failures are a reachability question.
Pull evidence
One item at a time, and nothing here tells you which one matters. Deciding what is worth looking at is most of the diagnosis.
0 of 9 inspected. You are not required to open all of them — a real investigation is judged on how few you needed.
What is your diagnosis?
Commit to one. Nothing below is shown until you do.
Guessing wrong and being told exactly why is the point of this page. Reading the answer first is not practice.