2.4 LB & Failover
System availability is uptime as a percentage. Load balancing distributes traffic; failover shifts it on failure. The arithmetic from nines to minutes is the worked example in this section.
2.4 System Availability: Load Balancing and Failover
Availability in nines
| Nines | Uptime % | Annual downtime |
|---|---|---|
| Two | 99.0 | ~3.65 days |
| Three | 99.9 | ~8.77 hours |
| Four | 99.99 | ~52.6 minutes |
| Five | 99.999 | ~5.26 minutes |
| Six | 99.9999 | ~31.5 seconds |
365 days x 24 hours x 60 minutes = 525,600 minutes per year. Every extra nine costs roughly 10x in redundancy.
Load balancing - types and choices
| Type | Routing key | Use |
|---|---|---|
| L4 (TCP) | IP + port | Cheap, fast, opaque |
| L7 (HTTP) | URL, cookie, header | Application-aware |
| DNS-based | Weighted records | Coarse-grained |
| Anycast | Same IP everywhere | CDNs, global services |
Algorithms: round-robin, least-connections, weighted, hash-based.
Failover - cold, warm, hot
| Mode | Spare state | Cutover time |
|---|---|---|
| Cold | Hardware off | Hours; cheapest |
| Warm | Booted, idle | Minutes |
| Hot | Running + receiving | Sub-second; costly |
Active-active vs active-passive
Worked example - bank online banking
| Item | Value |
|---|---|
| SLA target | 99.95 % availability |
| Budget | ~263 minutes (4 hr 23 min) of downtime per year |
| Architecture | L7 LB + active-active app tier + DB primary/standby |
| Failover | Sub-second for app, ~30 sec for DB promotion |
Common pitfalls
- DNS TTL stuck at hours; failover takes hours to propagate.
- 'Active-active' database without conflict resolution; data loss on
split-brain.
- Health checks too lenient; failing nodes keep getting traffic.
Mentor’s tip: Nines convert to minutes via 525,600. Every extra nine multiplies redundancy cost by ~10. L4 vs L7 vs DNS vs anycast - pick by what you actually need to route on. Failover speed is bounded by your state model, not your hardware.
Discussion
Loading…