CIO Brief: When Your Security Tool Causes the Outage
Retrospective: this article looks back at events from July 2024, written in 2026 with the benefit of hindsight.
The short version: On July 19, 2024, a faulty update to CrowdStrike's security software crashed about 8.5 million Windows computers worldwide. Airlines, hospitals and banks stopped working. It wasn't a cyberattack — the security tool itself caused the outage.
Why security tools are a resilience risk
Security software runs with the deepest access on every computer. That's how it stops attacks. It also means a bad update can disable every device at once. The same is true of other software deployed everywhere, such as operating system updates and management agents.
The business impact
- Company-wide outage in minutes.
- Slow manual recovery for encrypted devices.
- Lost revenue and customer disruption, with limited recourse against the vendor.
Questions to ask your team
- If every Windows computer and server failed to start tomorrow morning, how long would recovery take?
- Can we get encryption recovery keys quickly for thousands of devices?
- Do our security vendors let us stage updates rather than receive them everywhere at once?
- Do we have a way to communicate with employees if company laptops are down?
What good looks like
Staged updates where possible, recovery keys retrievable at scale, a tested mass-recovery runbook, and contracts that address vendor-caused outages.
The decision
Add "our security vendor causes an outage" to your next resilience exercise. It's realistic, and the answers often reveal recovery gaps that also matter for ransomware.
- The CrowdStrike Outage (July 19, 2024): 8.5 Million Windows Machines Down Incident Teardowns
- How to Recover BitLocker-Protected Azure VMs and Endpoints at Scale How-To & Hardening
- Endpoint Agent Update Risk Checklist How-To & Hardening