Engineering posts about Incident Management

Curated summaries and key learnings for engineers working with Incident Management.

Microsoft
5m

Try Azure SRE Agent with no always-on charges

The Azure SRE Agent offers a 30-day trial for new customers, allowing them to create and configure agents without incurring setup charges. The agent integrates with telemetry, source code, and...

Databricks
10m

How Databricks Uses AI to Accelerate Incident Investigation

The article discusses Databricks' implementation of AI SRE, an AI-powered debugging agent designed to streamline incident investigation across numerous microservices and Kubernetes clusters. It...

DigitalOcean
15m

Patching at Fleet Scale, Twice: How DigitalOcean Closed Januscape and the AMD Safe RET Issue Without Customer Impact

The article outlines DigitalOcean's response to two significant security vulnerabilities affecting their hypervisor fleet: Januscape and an AMD Safe RET issue. The first vulnerability required rapid...

Cloudflare
9m

Cloudflare DDoS Threat Report H1 2026: 1 Tbps attacks soar as DNS floods and geopolitical tensions drive a new wave

The Cloudflare DDoS Threat Report for the first half of 2026 highlights a significant increase in DDoS attack volume, with over 935 network-layer attacks exceeding 1 Tbps mitigated. The report...

Figma
20m

How we secure Figma’s internal systems with agents

The article discusses Figma's innovative approach to securing its internal systems through the development of an AI agent that enhances alert triage and forensic investigations. By leveraging a...

Cloudflare
9m

Cloudflare proudly joins the UK government's Cyber Resilience Pledge

Cloudflare's commitment to the UK's Cyber Resilience Pledge emphasizes the importance of cybersecurity governance and collective defense against cyber threats. The article outlines the organization's...

Cloudflare
25m

Build your own vulnerability harness

This article provides a comprehensive guide on constructing a model-agnostic vulnerability harness for enterprise codebases, emphasizing the importance of interchangeable AI models in enhancing...

Cloudflare
6m

Turning Cloudflare’s threat indicators into real-time WAF rules

The article presents a new integration of Cloudflare’s threat intelligence into Web Application Firewall (WAF) rules, allowing security teams to automate the blocking of high-risk IP addresses...

AWS
6m

Introducing the next generation of AWS Resilience Hub for generative AI-based SRE resilience journey

The article introduces the next generation of AWS Resilience Hub, emphasizing its enhancements for Site Reliability Engineers (SREs) in managing application resilience. Key features include a new...

Databricks
5m

How security teams can report cyber risk to boards

The article outlines the importance of translating cyber risk into financial terms to enable boards to make informed decisions regarding security investments. It emphasizes the need for coherent risk...

Duolingo
5m

Triage, ship, debug—all from Slack

Duolingo has developed an AI-powered Slack app that integrates with various tools such as GitHub, Jenkins, and AWS to enhance developer productivity and streamline incident management. The app...

Databricks
4m

Why AI Security Infrastructure is Now a CMO Priority

The article emphasizes the critical role of AI security infrastructure in modern enterprises, particularly highlighting the launch of Databricks Lakewatch, an innovative security information and...

Cloudflare
11m

When DNSSEC goes wrong: how we responded to the .de TLD outage

The article discusses the DNSSEC outage affecting the .de TLD on May 5, 2026, when DENIC published incorrect DNSSEC signatures, leading to widespread SERVFAIL responses from validating resolvers. It...

Cloudflare
11m

Code Orange: Fail Small is complete. The result is a stronger Cloudflare network

The article outlines the completion of Cloudflare's 'Code Orange: Fail Small' initiative, aimed at enhancing the resilience and reliability of its network infrastructure. Key improvements include the...

Databricks
4m

Alert Fatigue Is a Business Risk

The article highlights the critical issue of alert fatigue in enterprise security operations, where the overwhelming volume of alerts leads to significant risks as analysts struggle to prioritize and...

DigitalOcean
13m

From Incident Counting to SLIs: How DigitalOcean Rethought Availability

The article discusses DigitalOcean's transition from an incident-counting methodology to a more nuanced SLI-based approach for measuring availability. Initially, the company relied on a simplistic...

Meta (Facebook)
1m

Trust But Canary: Configuration Safety at Scale

In the Meta Tech Podcast episode featuring Pascal Hartig, the discussion revolves around the strategies employed by Meta's Configurations team to ensure safe configuration rollouts at scale. The...