On a Tuesday morning, a load balancer in our Klang Valley region failed over to its standby. The failover completed, and no customer lost data. It took eleven minutes. Our target for this class of failover is ninety seconds. This is the timeline, the root cause, and the changes we have made. We publish it because an incident reviewed only internally teaches one team a lesson, while an incident reviewed in public teaches everyone.
The timeline
The first alert fired at 09:04 MYT. The on-call engineer acknowledged it at 09:05, inside our fifteen-minute acknowledgement commitment. By 09:06 it was clear that the primary node was not recovering on its own, and the decision to fail over was taken. Traffic moved to the standby at 09:07. The standby then began failing health checks on the backends it had inherited, and the pool cycled through members for nine minutes before it settled at 09:15. Eleven minutes from first alert to stable service.
Nothing in this sequence happened during a change. There was no deployment in progress, no configuration change in the preceding twenty-four hours, and no maintenance activity on either node. The primary stopped responding to its peers, which is the condition failover exists to handle, and the cluster did what it was designed to do in response. The failure was in the state the standby had been left in, not in the failover path itself.
Monitoring saw the incident before any customer did. The alert came from the control plane rather than from a support ticket. Within forty seconds of traffic moving, the backend pool showed check failures rising from zero to more than half the pool, while the backends themselves continued to answer requests that reached them. That split, healthy servers marked unhealthy by the node in front of them, was the signal that told us where to look.
The root cause
The standby node had a healthy heartbeat but a stale health-check configuration. Heartbeat checks told us the node was alive; they said nothing about whether it was configured like its peer. The settings that mattered were the check interval, the response timeout, the number of consecutive failures before a backend is removed, and the drain timeout applied when a backend leaves the pool. Those settings lived in a local file on each node rather than in the replicated state.
The two nodes had drifted. The primary had been tuned over months as backends were added and their response profiles changed, and its settings reflected that work. The standby still carried the values it was built with. Its checks were stricter and its timeouts shorter than the backends could satisfy while carrying a full share of traffic, so when traffic moved, the standby marked healthy backends as down and removed them from the pool.
The failure was not in the failover mechanism. The mechanism detected that the primary had stopped responding and moved traffic, exactly as designed. The failure was in the assumption that a healthy heartbeat implies an identical configuration. That assumption had never been tested under production load, because the standby had never carried production traffic for more than a few minutes during a maintenance window.
- Check interval and response timeout
- Consecutive failure threshold before removal
- Drain timeout applied at failover
- Backend weights and pool membership
Why the pool took nine minutes to settle
Once the standby was in front of live traffic, it began removing backends that were in fact serving requests normally. Traffic was retried across the remaining pool, and because sessions were pinned, some clients were sent to members that were themselves being marked down. The pool thrashed: members left, checks passed again, members returned, checks failed again. Each cycle is fast, but the threshold that finally stabilised the pool was reached only after nine minutes.
Two mechanisms extended the recovery. The first was connection draining. The drain timeout on the standby was shorter than on the primary, so connections that would have been allowed to finish were closed and retried instead. The second was the passive check path, which marks a backend down on the basis of live request failures and can oscillate when a backend is healthy but slow to answer under load. Neither mechanism was broken; both were configured differently from the node they replaced.
The failover itself was fast. Measured from the moment the standby accepted traffic to the moment the pool was stable, the mechanism took under two seconds. Nine of the eleven minutes were spent waiting for check thresholds to settle, and the last minutes were spent verifying that service had returned to normal before standing the incident down. The mechanism we had tested was quick; the configuration it inherited was not.
Impact
The incident fell inside a maintenance window that had been announced seven days in advance, so no customer workload was interrupted and no data was lost. Requests in flight during the switch either completed on the previous node or were retried successfully by clients. We saw no sustained rise in error rate above our internal threshold, and no customer raised a ticket during the window. The customer-facing availability commitment for the month was met.
What we changed
- Health-check configuration is now part of the replicated state
- Standby nodes run synthetic traffic against a shadow backend
- Failover drills now run monthly, during MYT business hours
- Drill results are published in the monthly service review
- Load balancer configuration changes now need a second reviewer
Health-check configuration is now replicated. When an engineer changes the check interval, the response timeout, the failure threshold or the drain timeout on one node, the change is written to the replicated state and applied to every node in the pair. A standby can no longer hold a configuration that differs from its primary, because the configuration is no longer stored only on the node. The review checklist for any load balancer change now includes those four parameters.
Standby nodes now run synthetic traffic against a shadow backend continuously, not only during drills. The probe exercises the same check path that real traffic would, from the same node, with the same configuration. A standby that cannot serve is now detected by a monitor rather than by a customer request. We also keep a small share of production traffic on the standby outside maintenance windows, so the configuration is exercised rather than assumed.
Failover drills now run monthly, during MYT business hours, and they are scheduled rather than improvised. A drill fails if the pool does not settle inside the ninety-second target, and a failed drill creates a work item with the same priority as a production defect. Drill results, including the measured time to stable service, are published in the monthly service review alongside availability and incident counts.
Changes to the load balancer now follow the same review path as any other production change. A second engineer reviews the diff, the change is applied to staging first, and the health-check parameters are part of the review checklist rather than an implementation detail. A change that alters check thresholds cannot be merged by its author alone, regardless of how small the diff appears.
We are also reducing reliance on any single pair of nodes. Capacity in the region is spread across availability zones, and the load balancer tier is being extended so that a failure in one zone is absorbed by the others rather than by a standby in the same zone. That work was already funded and scheduled; this incident changed its sequencing, not its scope.
How we will know the fix holds
We track four numbers for this tier: time from first alert to traffic on the standby, time from traffic moving to a stable pool, the rate of check flapping during a failover, and the number of configuration differences detected between paired nodes. The first two have targets, the third should be zero, and the fourth should be zero. If a drill produces a difference between paired nodes, we treat it as a defect and fix the replication path rather than the node.
The customer-facing SLA did not change, and it was met. The eleven minutes fell inside the maintenance window, so no customer was affected. The margin was thinner than it should have been, and we would rather learn that from a drill than from an incident that lands outside a window. The commitments customers rely on are unchanged: 99.95% availability, acknowledgement of security incidents within fifteen minutes, support during MYT business hours, and a written review within five business days of any incident that affects customers.
A standby that has never carried traffic is a hypothesis, not a redundancy.
We publish this because the alternative is that the same lesson is learned privately and repeated somewhere else. The changes above are in production, the drill is on the calendar, and the numbers are in the monthly review. If the next failover is not faster, the review will say so, in the same place this postmortem appears.



