Region Failure Readiness Checklist for AWS Workloads
Retrospective: this article looks back at events from October 2025, written in 2026 with the benefit of hindsight.
Use this checklist to check whether an AWS workload is ready for a regional failure.
Business requirements
- Recovery time objective (RTO) and recovery point objective (RPO) agreed with the business.
- DR strategy chosen and funded (backup/restore, pilot light, warm standby, active-active).
Dependencies
- AWS services the workload depends on documented, including regional and global dependencies.
- Third-party SaaS dependencies documented, with their regional exposure.
- DNS, identity and certificate dependencies identified.
Data
- Replication configured for databases and storage.
- Replication lag monitored.
- Backups copied to another region and account.
Infrastructure
- Infrastructure as code can deploy the workload in the secondary region.
- Capacity pre-provisioned for warm standby or active-active.
- Secrets and KMS keys available in the secondary region.
Failover mechanism
- Route 53 health checks and failover records configured.
- Failover doesn't require control-plane changes in the failed region.
- Runbook documented and accessible outside AWS.
Operations
- Monitoring and alerting work in the secondary region.
- Status communication plan for customers.
- Break-glass access to AWS not dependent on a single region or SSO path.
Testing
- Failover tested within the past 12 months.
- Actual RTO and RPO measured.
- Gaps tracked to closure.
- The AWS us-east-1 Outage of October 2025: DNS, DynamoDB and a Day of Downtime Incident Teardowns
- How to Reduce Your Dependence on a Single AWS Region's Control Plane How-To & Hardening
- CIO Brief: Cloud Concentration Risk Is Back on the Board Agenda CIO Briefings