Engineering posts about Failover
Curated summaries and key learnings for engineers working with Failover.
From Silos to Service Topology: Why Netflix Built a Real-Time Service Map
The article outlines Netflix's development of a real-time service topology map to improve observability and troubleshooting in its microservices architecture. It highlights the challenges faced by...
Sitar-agent: Building a reliable dynamic configuration sidecar at scale
The article discusses the development of Sitar-agent, a Kubernetes sidecar designed to ensure reliable dynamic configuration delivery at scale for Airbnb's services. It outlines the configuration...
Lights Out, Systems On: Validating Instant Power Loss Readiness
The article introduces the Instantaneous PowerLoss Storm, a testing paradigm developed by Meta to prepare data centers for zero-notice power loss scenarios. It outlines the strategies implemented to...
How the lakebase architecture stays resilient to cloud failures
The article discusses the challenges faced by cloud infrastructure due to increased demand for control-plane operations and the need for high availability in database management. It outlines how the...