Engineering posts about Incident Management
Curated summaries and key learnings for engineers working with Incident Management.
Try Azure SRE Agent with no always-on charges
The Azure SRE Agent offers a 30-day trial for new customers, allowing them to create and configure agents without incurring setup charges. The agent integrates with telemetry, source code, and...
How Databricks Uses AI to Accelerate Incident Investigation
The article discusses Databricks' implementation of AI SRE, an AI-powered debugging agent designed to streamline incident investigation across numerous microservices and Kubernetes clusters. It...
Patching at Fleet Scale, Twice: How DigitalOcean Closed Januscape and the AMD Safe RET Issue Without Customer Impact
The article outlines DigitalOcean's response to two significant security vulnerabilities affecting their hypervisor fleet: Januscape and an AMD Safe RET issue. The first vulnerability required rapid...
Cloudflare DDoS Threat Report H1 2026: 1 Tbps attacks soar as DNS floods and geopolitical tensions drive a new wave
The Cloudflare DDoS Threat Report for the first half of 2026 highlights a significant increase in DDoS attack volume, with over 935 network-layer attacks exceeding 1 Tbps mitigated. The report...
How we secure Figma’s internal systems with agents
The article discusses Figma's innovative approach to securing its internal systems through the development of an AI agent that enhances alert triage and forensic investigations. By leveraging a...
Cloudflare proudly joins the UK government's Cyber Resilience Pledge
Cloudflare's commitment to the UK's Cyber Resilience Pledge emphasizes the importance of cybersecurity governance and collective defense against cyber threats. The article outlines the organization's...
Build your own vulnerability harness
This article provides a comprehensive guide on constructing a model-agnostic vulnerability harness for enterprise codebases, emphasizing the importance of interchangeable AI models in enhancing...
Turning Cloudflare’s threat indicators into real-time WAF rules
The article presents a new integration of Cloudflare’s threat intelligence into Web Application Firewall (WAF) rules, allowing security teams to automate the blocking of high-risk IP addresses...
Introducing the next generation of AWS Resilience Hub for generative AI-based SRE resilience journey
The article introduces the next generation of AWS Resilience Hub, emphasizing its enhancements for Site Reliability Engineers (SREs) in managing application resilience. Key features include a new...
How security teams can report cyber risk to boards
The article outlines the importance of translating cyber risk into financial terms to enable boards to make informed decisions regarding security investments. It emphasizes the need for coherent risk...
Triage, ship, debug—all from Slack
Duolingo has developed an AI-powered Slack app that integrates with various tools such as GitHub, Jenkins, and AWS to enhance developer productivity and streamline incident management. The app...
Why AI Security Infrastructure is Now a CMO Priority
The article emphasizes the critical role of AI security infrastructure in modern enterprises, particularly highlighting the launch of Databricks Lakewatch, an innovative security information and...
When DNSSEC goes wrong: how we responded to the .de TLD outage
The article discusses the DNSSEC outage affecting the .de TLD on May 5, 2026, when DENIC published incorrect DNSSEC signatures, leading to widespread SERVFAIL responses from validating resolvers. It...
Code Orange: Fail Small is complete. The result is a stronger Cloudflare network
The article outlines the completion of Cloudflare's 'Code Orange: Fail Small' initiative, aimed at enhancing the resilience and reliability of its network infrastructure. Key improvements include the...
Alert Fatigue Is a Business Risk
The article highlights the critical issue of alert fatigue in enterprise security operations, where the overwhelming volume of alerts leads to significant risks as analysts struggle to prioritize and...
From Incident Counting to SLIs: How DigitalOcean Rethought Availability
The article discusses DigitalOcean's transition from an incident-counting methodology to a more nuanced SLI-based approach for measuring availability. Initially, the company relied on a simplistic...
Trust But Canary: Configuration Safety at Scale
In the Meta Tech Podcast episode featuring Pascal Hartig, the discussion revolves around the strategies employed by Meta's Configurations team to ensure safe configuration rollouts at scale. The...