Home → DevOps → Article

Incident.io Eliminates Single Point of Failure With Dual Message Broker

TL;DR

Engineering team built dynamic load balancer across Pub/Sub and NATS, achieving 99.99% availability with zero customer impact during failover testing.

Key Points

  • Processing ~240 million messages daily across 800+ topics with 1000+ subscriptions
  • Implemented delay-based MaxWeight scheduler inspired by queueing theory (OCF algorithm)
  • Successfully disabled Google Cloud Pub/Sub in production with zero dropped messages or customer-facing errors
  • Active-active 50/50 load balancing with automated circuit breaker failover in <30 seconds

Why It Matters

This demonstrates a sophisticated approach to eliminating vendor lock-in and single points of failure in event-driven architectures at scale. The delay-based scheduling algorithm and chaos testing methodology provide a reusable pattern for teams managing mission-critical messaging infrastructure with strict SLA requirements.
Read the technical deep-dive

Source: incident.io