The Failover That Lost Ninety Seconds of Orders
Read what each party saw, commit to a cause, and only then find out which of them was right. The root cause, why the obvious reading was wrong, and the fixes are all held back until you have answered.
The symptom
What was reported, before anyone knew what was happening.
What each party saw, and what each concluded
The evidence, in the form it actually arrived in — several parties, several partial views, several confident conclusions.
| Who | What they could actually see | What they concluded | Verdict |
|---|---|---|---|
| Failover controller | The primary unreachable for 90 seconds; the secondary replica healthy and accepting connections. | Promote the secondary. Recovery time objective met. | ✕ wrong |
| Secondary replica | Its own log ending at a position it had no way to compare against the primary’s. | I am the primary now; accept writes from this position. | ✕ wrong |
| Payment provider | Charges captured and settled before the outage. | Nothing — the money moved and stayed moved. | ✓ right |
| Incident commander | RTO of 4 minutes against a 15-minute target, no errors after failover. | A textbook failover; close the incident. | ✕ wrong |
3 of 4 parties reasoned correctly from what they could see and still reached the wrong conclusion. Nobody in this table is careless. Each one acted on complete-looking local information, and the information was local. That gap — between what a node can observe and what is true — is the whole domain, and one of these readings will usually be yours.
Commit before you read on
The failover met its recovery time objective and the service was healthy afterwards. Commit before reading on: name the objective nobody checked, and say what the secondary would have needed to know to make promotion safe.
Write it down even if you are unsure. An unwritten guess quietly becomes “that is what I thought” the moment you read the answer.