AWS Training
Modules Listen All tracks

← Design Resilient Architectures

Starts this lesson and continues through 14 more to the end of certification prep.

Disaster recovery — the four strategies and the data plane rule

Why this lesson closes the module

Task statement 2.2 names them explicitly: "Disaster recovery (DR) strategies (for example, backup and restore, pilot light, warm standby, active-active failover, recovery point objective [RPO], recovery time objective [RTO])", plus the skill "Selecting an appropriate DR strategy to meet business requirements."

That is the most prescriptive item in the whole exam guide. It names the four strategies and the two metrics that choose between them. So the exam question writes itself: here is an RTO and an RPO, pick a strategy.

The two metrics

The whitepaper's framing of the backup cadence: "How often you run your backup will determine your achievable recovery point (which should align to meet your RPO)."

The four strategies, in cost order

From Disaster recovery options in the cloud, verified 2026-09-22:

"Disaster recovery strategies available to you within AWS can be broadly categorized into four approaches, ranging from the low cost and low complexity of making backups to more complex strategies using multiple active Regions. Active/passive strategies use an active site (such as an AWS Region) to host the workload and serve traffic. The passive site … is used for recovery. The passive site does not actively serve traffic until a failover event is triggered."

   COST / COMPLEXITY  ──────────────────────────────────────────▶
   RTO / RPO          ◀──────────────────────────────────────────

   Backup & restore  →  Pilot light  →  Warm standby  →  Multi-site active/active
   ├ data backed up     ├ data replicating   ├ data replicating    ├ live in every Region
   ├ nothing running    ├ core infra ON      ├ everything ON,      ├ all Regions serve
   └ redeploy all       └ servers OFF        │  scaled down        └ "no such thing as
                                             └ scale up only          failover"
        active/passive ──────────────────────────────┘

1. Backup and restore

"Backup and restore is a suitable approach for mitigating against data loss or corruption. This approach can also be used to mitigate against a regional disaster by replicating data to other AWS Regions… In addition to data, you must redeploy the infrastructure, configuration, and application code in the recovery Region."

⚠️ The infrastructure-as-code dependency is the point:

"To enable infrastructure to be redeployed quickly without errors, you should always deploy using infrastructure as code (IaC) using services such as AWS CloudFormation or the AWS Cloud Development Kit (AWS CDK). Without IaC, it may be complex to restore workloads in the recovery Region, which will lead to increased recovery times and possibly exceed your RTO."

Back up "code and configuration, including Amazon Machine Images (AMIs)", not just user data.

When it's enough, verbatim: "For a disaster event based on disruption or loss of one physical data center for a well-architected, highly available workload, you may only require a backup and restore approach." Beyond that — "If your definition of a disaster goes beyond the disruption or loss of a physical data center to that of a Region" — you need one of the other three.

AWS Backup supports "copying backups across Regions", and "The cross-account backup capability helps protect from disaster events that include insider threats or account compromise" (SAA1 lesson 6).

⚠️ Restore is not automatic. "AWS Backup offers restore capability, but does not currently enable scheduled or automatic restoration." And the reason to rehearse it: "Implementing a scheduled periodic data restore is a good idea as data restore from backup is a control plane operation. If this operation was not available during a disaster, you would still have operable data stores created from a recent backup."

S3 versioning as a human-error control: "Object versioning protects your data in S3 from the consequences of deletion or modification actions by retaining the original version… can be a useful mitigation for human-error type disasters."

And a subtle protection worth knowing: *"by default, when an object is deleted in the source bucket, Amazon S3 adds a delete marker in the source bucket only. This approach protects data in the DR Region from malicious deletions in the source Region."*

2. Pilot light

"With the pilot light approach, you replicate your data from one Region to another and provision a copy of your core workload infrastructure. Resources required to support data replication and backup, such as databases and object storage, are always on. Other elements, such as application servers, are loaded with application code and configurations, but are "switched off" and are only used during testing or when disaster recovery failover is invoked."

AWS's own best-practice note is a nuance worth having: "A best practice for 'switched off' is to not deploy the resource, and then create the configuration and capabilities to deploy it ('switch on') when needed."

Continuous replication is what buys the low RPO. The services AWS names:

⚠️ Aurora global database is the standout number:

"Global database uses dedicated infrastructure that leaves your databases entirely available to serve your application, and can replicate to the secondary Region with typical latency of under a second (and within an AWS Region is much less than 100 milliseconds). With Amazon Aurora global database, if your primary Region suffers a performance degradation or outage, you can promote one of the secondary regions to take read/write responsibilities in less than one minute even in the event of a complete regional outage."

Compare plain RDS: "you must promote an RDS read replica to become the primary instance. For DB instances other than Aurora, the process takes a few minutes to complete and rebooting is part of the process."

⚠️ Continuous replication is not a backup. "Continuous replication of data has the advantage of being the shortest time (near zero) to back up your data, but may not protect against disaster events such as data corruption or malicious attack (such as unauthorized data deletion) as well as point-in-time backups." Replication faithfully copies your corruption. You need both.

3. Warm standby

"The warm standby approach involves ensuring that there is a scaled down, but fully functional, copy of your production environment in another Region. This approach extends the pilot light concept and decreases the time to recovery because your workload is always-on in another Region."

⚠️ The distinction AWS knows you'll struggle with — it says so in a boxed note:

"The difference between pilot light and warm standby can sometimes be difficult to understand. Both include an environment in your DR Region with copies of your primary Region assets. The distinction is that pilot light cannot process requests without additional action taken first, whereas warm standby can handle traffic (at reduced capacity levels) immediately. The pilot light approach requires you to 'turn on' servers, possibly deploy additional (non-core) infrastructure, and scale up, whereas warm standby only requires you to scale up."

One-line test: can the DR Region serve a request right now, without you doing anything? No → pilot light. Yes, but slowly → warm standby.

4. Multi-site active/active — and hot standby

"You can run your workload simultaneously in multiple Regions as part of a multi-site active/active or hot standby active/passive strategy. Multi-site active/active serves traffic from all regions to which it is deployed, whereas hot standby serves traffic only from a single region, and the other Region(s) are only used for disaster recovery."

"This approach is the most complex and costly approach to disaster recovery, but it can reduce your recovery time to near zero for most disasters… With multi-site active/active, because the workload is running in more than one Region, there is no such thing as failover in this scenario."

AWS's pragmatic aside is worth repeating in an interview: "Most customers find that if they are going to stand up a full environment in the second Region, it makes sense to use it active/active. Alternatively, if you do not want to use both Regions to handle user traffic, then Warm Standby offers a more economical and operationally less complex approach."

⚠️ Even active/active doesn't give you zero RPO for a data disaster: "recovery times for a data disaster involving data corruption, deletion, or obfuscation will always be greater than zero and the recovery point will always be at some point before the disaster was discovered."

The write strategies for active/active — the genuinely hard part, and examinable:

Strategy What it does AWS's named fit
Write global "routes all writes to a single Region. In case of failure of that Region, another Region would be promoted" Aurora global database — "supports write forwarding"
Write local "routes writes to the closest Region (just like reads)" DynamoDB global tables — "use a last writer wins reconciliation between concurrent updates"
Write partitioned "assigns writes to a specific Region based on a partition key (like user ID) to avoid write conflicts" S3 bi-directional replication (two Regions), with "replica modification sync" enabled

"It is common to design user reads to be served from the Region closest to them, known as read local."

The rule that decides whether your failover works

This is the most important idea in the lesson and it appears near the top of the whitepaper:

"When choosing your strategy, and the AWS resources to implement it, keep in mind that within AWS, we commonly divide services into the data plane and the control plane. The data plane is responsible for delivering real-time service while control planes are used to configure the environment. For maximum resiliency, you should use only data plane operations as part of your failover operation. This is because the data planes typically have higher availability design goals than the control planes."

Apply it and a lot of design choices invert:

Failover mechanism Plane Verdict
Route 53 health checks driving DNS failover data "a highly reliable operation done on the data plane"
Amazon Application Recovery Controller (ARC) routing controls data "you can script failover using this highly available, data plane API"
Changing Route 53 weights to shift traffic control "be aware this is a control plane operation and therefore not as resilient"
Global Accelerator traffic dials control "note this is a control plane operation"
Auto Scaling to reach production capacity control "taking a dependency on it will lower the resiliency of your overall recovery strategy"
Restoring from AWS Backup control rehearse it; keep restored stores standing by

⚠️ This is why "hot standby" exists as a concept. AWS: "You can choose to provision sufficient capacity such that the recovery Region can handle the full production load as deployed. This statically stable configuration is called hot standby… Or you may choose to provision fewer resources which will cost less, but take a dependency on Auto Scaling."

Static stability: a system that needs to do nothing — call no API, provision nothing, scale nothing — in order to survive the failure. It costs more. It works when the control plane doesn't.

Automatic or manual failover?

AWS is notably cautious, and the exam reflects it:

"This failover operation can be initiated either automatically or manually. Automatically initiated failover based on health checks or alarms should be used with caution. Even using the best practices discussed here, recovery time and recovery point will be greater than zero, incurring some loss of availability and data. If you fail over when you don't need to (false alarm), then you incur those losses. Manually initiated failover is therefore often used. In this case, you should still automate the steps for failover, so that the manual initiation is like the push of a button."

Manual trigger, automated steps. That phrasing is worth memorising — it's the correct answer shape to "should DR failover be automatic?"

Traffic management at failover

Option Verbatim Note
Route 53 health checks "you can configure automatically initiated DNS failover to ensure traffic is sent only to healthy endpoints" data plane; subject to DNS caching
Application Recovery Controller "create Route 53 health checks that do not actually check health, but instead act as on/off switches that you have full control over" data plane; the manual-trigger answer
Global Accelerator "Using AnyCast IP, you can associate multiple endpoints in one or more AWS Regions with the same static public IP address"; "avoids caching issues that can occur with DNS systems" lower latency, no DNS caching
CloudFront origin failover "if a given request to the primary endpoint fails, CloudFront routes the request to the secondary endpoint" ⚠️ "all subsequent requests still go to the primary endpoint, and failover is done per each request"

⚠️ That CloudFront note is a real distinction: origin failover is per request, not a sticky switch-over.

Testing — the part everyone skips

"It is critical to regularly assess and test your disaster recovery strategy so that you have confidence in invoking it, should it become necessary. Use AWS Resilience Hub to continuously validate and track the resilience of your AWS workloads, including whether you are likely to meet your RTO and RPO targets."

And for active/active, testing changes shape: "there is no such thing as failover in this scenario. Disaster recovery testing in this case would focus on how the workload reacts to loss of a Region: Is traffic routed away from the failed Region? Can the other Region(s) handle all the traffic?"

That last question is the one that catches people — two Regions each running at 60% cannot absorb each other.

This is SAA0 lesson 1's "improve through game days" principle, applied. Documenting the runbook is not testing it.

Mapping a requirement to a strategy

The stem says Strategy
"RPO 24 hours, RTO 24 hours, lowest cost" Backup and restore
"RPO minutes, RTO tens of minutes, keep cost down" Pilot light
"RPO seconds, RTO minutes, can't wait to deploy" Warm standby
"RTO and RPO near zero, cost is secondary" Multi-site active/active
"serve users from the nearest Region and survive losing one" active/active, read local
"sub-second cross-Region replication, promote in under a minute" Aurora global database
"writes accepted in every Region" DynamoDB global tables (last writer wins)
"protect against ransomware / malicious deletion" backups + versioning + Object Lock — not replication alone
"failover must work even if the control plane is degraded" hot standby / static stability, ARC routing controls

Check yourself

  1. In one sentence, what separates pilot light from warm standby?
  2. Why does AWS warn against using Route 53 weights to shift traffic during a failover?
  3. Your DR plan depends on Auto Scaling to reach production capacity. What's the trade-off, and what's the alternative called?
  4. Should DR failover be automatic?
  5. You replicate continuously to another Region. Are you protected against an engineer deleting a table?
Answers
  1. Pilot light cannot serve a request until you act; warm standby can, at reduced capacity. AWS: "pilot light cannot process requests without additional action taken first, whereas warm standby can handle traffic (at reduced capacity levels) immediately."
  2. Because it's a control plane operation — "not as resilient as the data plane approach using Amazon Application Recovery Controller." Control planes have lower availability design goals, and a Regional event is exactly when they're stressed.
  3. Trade-off: "Because Auto Scaling is a control plane activity, taking a dependency on it will lower the resiliency of your overall recovery strategy." The alternative — provisioning full capacity in advance — is hot standby, and the property is static stability.
  4. Usually manual trigger with automated steps. "Automatically initiated failover … should be used with caution… If you fail over when you don't need to (false alarm), then you incur those losses… you should still automate the steps for failover, so that the manual initiation is like the push of a button."
  5. No. "Continuous replication … may not protect against disaster events such as data corruption or malicious attack (such as unauthorized data deletion) as well as point-in-time backups." Replication copies the deletion. You need point-in-time backups, versioning, and — from SAA1 lesson 6 — Object Lock, because permissions are not immutability.

Teaching this section

← PreviousHigh availability — surviving an instance, an AZ, a dependencyFinished →Cheat sheet, lab & quiz