Engineering posts about Alerting
Curated summaries and key learnings for engineers working with Alerting.
Scaling Security Alert Triage With Specialized Agents on Databricks
The article discusses how Databricks has implemented a specialized agent-based system for triaging security alerts, particularly focusing on low-severity alerts that are often overlooked. By...
Using observability data to prevent incidents
The article emphasizes the importance of using observability data to transition from reactive incident response to proactive reliability intelligence. It outlines how engineering teams can leverage...
Monitoring reliably at scale
The article outlines the challenges of maintaining reliable observability in systems that are heavily dependent on shared infrastructure, such as Kubernetes and service meshes. It highlights the...
Trust But Canary: Configuration Safety at Scale
In the Meta Tech Podcast episode featuring Pascal Hartig, the discussion revolves around the strategies employed by Meta's Configurations team to ensure safe configuration rollouts at scale. The...