Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Starts this lesson and continues through 15 more to the end of certification prep.
Task statement 2.2 is "Design highly available and/or fault-tolerant architectures", and its skills list is explicit about the unit of failure: "Determining the AWS services required to provide a highly available and/or fault-tolerant architecture across AWS Regions or Availability Zones" and "Implementing designs to mitigate single points of failure".
So the method is mechanical. Walk the request path, and at each hop ask: if this one thing dies, what happens? Every answer that is "the system goes down" is a single point of failure, and every one has a standard fix.
client → DNS → load balancer → compute → database → dependency
│ │ │ │ │
│ │ │ │ └ connection pool exhaustion → RDS Proxy
│ │ │ └ instance loss → Multi-AZ
│ │ └ instance loss → Auto Scaling across AZs
│ └ AZ loss → subnets in ≥2 AZs
└ endpoint loss → Route 53 health checks + failover
The task statement says "highly available and/or fault-tolerant", and they are not synonyms:
| Highly available | Fault tolerant | |
|---|---|---|
| Promise | recovers quickly | keeps running with no interruption |
| Example | Multi-AZ RDS failover (brief outage) | multiple active instances behind a load balancer |
| Cost | lower | higher — you pay for idle redundancy |
⚠️ Multi-AZ RDS is highly available, not fault tolerant. Failover is automatic, but there is a gap. If a stem demands "zero downtime, no dropped connections", a standby that has to be promoted doesn't strictly meet it — that's an Aurora cluster with multiple readers, or an active-active design. The exam usually isn't that pedantic, but knowing the distinction stops you over-reading a question.
Covered in lesson 2, restated as the resilience answer: an Auto Scaling group "automatically monitors the health and availability of your instances using EC2 health checks and replaces terminated or impaired instances to maintain your desired capacity", plus custom health checks where the EC2 check isn't enough.
Min = max = desired still self-heals. That's the cheapest fix for "an instance died and nobody noticed".
From SAA0 lesson 4, verbatim: "A subnet must reside within a single Availability Zone." Therefore
every AZ-resilient design is a multi-subnet design:
⚠️ The standby serves no reads. "A Multi-AZ standby cannot serve read requests. Multi-AZ deployments are designed to provide enhanced database availability and durability, rather than read scaling benefits." This is the highest-frequency RDS misconception and it recurs across Domains 2, 3 and 4.
Read replicas are the other tool: "make it easier to take advantage of supported engines' built-in replication functionality to elastically scale out beyond the capacity constraints of a single DB instance for read-heavy database workloads." Asynchronous, manual promotion — read scaling, not availability.
| Requirement | Answer |
|---|---|
| "survive the loss of an AZ" | Multi-AZ |
| "reads are overwhelming the primary" | read replicas |
| both | both — they're different features |
| "promote a copy in another Region" | cross-Region read replica, or Aurora global database |
That's lesson 6. But note the boundary now: everything above is within one Region. An Auto Scaling group does not span Regions. Multi-AZ does not span Regions. If the stem says "Regional outage", none of the controls in this section are sufficient.
From Active-active and active-passive failover, verified 2026-09-22:
"You configure active-active failover using any routing policy (or combination of routing policies) other than failover, and you configure active-passive failover using the failover routing policy."
⚠️ That's a precise, examinable statement and it's counterintuitive. The failover routing policy gives you active-passive. Active-active comes from weighted, latency, geolocation or multivalue — anything but failover.
Active-active, verbatim:
"Use this failover configuration when you want all of your resources to be available the majority of the time. When a resource becomes unavailable, Route 53 can detect that it's unhealthy and stop including it when responding to queries. … Route 53 can respond to a DNS query using any healthy record."
Active-passive, verbatim:
"Use an active-passive failover configuration when you want a primary resource or group of resources to be available the majority of the time and you want a secondary resource or group of resources to be on standby… When responding to queries, Route 53 includes only the healthy primary resources. If all the primary resources are unhealthy, Route 53 begins to include only the healthy secondary resources."
Note "If all the primary resources are unhealthy" — with multiple primaries, Route 53 "considers the primary failover record to be healthy as long as at least one of the associated resources is healthy." One surviving primary keeps all traffic on the primary side.
The alias-record rule, verbatim, and it's a real configuration trap:
"If you're routing traffic to any AWS resources that you can create alias records for, don't create health checks for those resources. When you create the alias records, you set Evaluate Target Health to Yes instead."
⚠️ So for an ALB, you do not attach a Route 53 health check — you set Evaluate Target Health. The load balancer already knows whether its targets are healthy; a separate health check duplicates it and can disagree with it.
The weighted-records caveat is worth knowing because it describes a self-inflicted outage:
"All the records with nonzero weights must be unhealthy before Route 53 starts to respond to DNS queries using records that have weights of zero. This can make your web application or website unreliable if the last healthy resource, such as a web server, can't handle all the traffic when other resources are unavailable."
That is: your last healthy server gets 100% of traffic and falls over. Cascading failure by design.
Task 2.2 lists "Proxy concepts (for example, Amazon RDS Proxy)". From Amazon RDS Proxy, verified 2026-09-22:
"By using Amazon RDS Proxy, you can allow your applications to pool and share database connections to improve their ability to scale. RDS Proxy makes applications more resilient to database failures by automatically connecting to a standby DB instance while preserving application connections."
That second sentence is the Domain 2 reason it exists: it preserves application connections across a failover. Without a proxy, a Multi-AZ failover drops every open connection and the application has to reconnect and retry.
The scaling problem it solves, verbatim:
"Using RDS Proxy, you can handle unpredictable surges in database traffic. Otherwise, these surges might cause issues due to oversubscribing connections or new connections being created at a fast rate. RDS Proxy establishes a database connection pool and reuses connections in this pool. This approach avoids the memory and CPU overhead of opening a new database connection each time."
And the load-shedding behaviour, which is a resilience pattern in itself:
"RDS Proxy queues or throttles application connections that can't be served immediately from the connection pool. Although latencies might increase, your application can continue to scale without abruptly failing or overwhelming the database. If connection requests exceed the limits you specify, RDS Proxy rejects application connections (that is, it sheds load). At the same time, it maintains predictable performance for the load that RDS can serve."
⚠️ Lambda + RDS is the classic RDS Proxy question. Lambda scales by adding execution environments (lesson 4), each opening its own connection — so a traffic spike can exhaust the database's connection limit. RDS Proxy pools them. If a stem pairs "serverless" with "too many database connections", that's the answer.
Limits worth knowing, verbatim: "Each AWS account ID is limited to 20 proxies" (adjustable via Service Quotas); "Each proxy can have up to 200 associated Secrets Manager secrets"; "For RDS DB instances in replication configurations, you can associate a proxy only with the writer DB instance, not a read replica"; and "Your RDS Proxy must be in the same virtual private cloud (VPC) as the database. The proxy can't be publicly accessible."
It also "can enforce AWS Identity and Access Management (IAM) authentication for clients connecting to the proxy", connecting onward with IAM auth or Secrets Manager credentials — the Domain 1 principle of eliminating long-lived credentials, applied.
Task 2.2 lists "Service quotas and throttling (for example, how to configure the service quotas for a workload in a standby environment)". That parenthetical is doing real work.
The DR whitepaper makes it concrete:
"Ensure that service quotas in your DR Region are set high enough so as to not limit you from scaling up to production capacity."
⚠️ A quota is a single point of failure you cannot see. Your DR Region looks configured, the failover runs, Auto Scaling tries to launch 200 instances, and the Region's quota is 20. This is an exam answer and a real incident. The fix is to raise quotas in the standby Region in advance and verify them.
Task 2.2 lists "Workload visibility (for example, AWS X-Ray)". From What is AWS X-Ray?, verified 2026-09-22:
"AWS X-Ray is a service that collects data about requests that your application serves, and provides tools that you can use to view, filter, and gain insights into that data to identify issues and opportunities for optimization. For any traced request to your application, you can see detailed information not only about the request and response, but also about calls that your application makes to downstream AWS resources, microservices, databases, and web APIs."
The trace map is the exam-relevant artefact:
"X-Ray uses trace data from the AWS resources that power your cloud applications to generate a detailed trace map. The trace map shows the client, your front-end service, and backend services that your front-end service calls to process requests and persist data. Use the trace map to identify bottlenecks, latency spikes, and other issues."
Mechanics worth a line: "each client SDK sends JSON segment documents to a daemon process listening for UDP traffic. The X-Ray daemon buffers segments in a queue and uploads them to X-Ray in batches." And it's "included on AWS Elastic Beanstalk and AWS Lambda platforms."
The discriminator against CloudWatch: CloudWatch tells you a service is slow. X-Ray tells you which downstream call is slow, for a specific request, across service boundaries. If a stem says "a microservices application is slow and we don't know which component", that's X-Ray — a metric per service can't answer it.
Given any architecture in an exam stem, walk it:
| Component | If it dies | Fix |
|---|---|---|
| Single EC2 instance | outage | Auto Scaling group, ≥2 AZs, behind an ELB |
| Single AZ | outage | subnets in ≥2 AZs, everywhere |
| Single-AZ RDS | outage | Multi-AZ |
| Overloaded primary DB | degradation | read replicas |
| Connection storm from Lambda | DB refuses connections | RDS Proxy |
| NAT gateway in one AZ | that AZ loses egress | NAT per AZ (see ⚠️ below) |
| One endpoint, no health check | traffic to a dead host | Route 53 health checks + failover |
| Quota in the standby Region | failover can't scale | raise quotas in advance |
| Unknown slow component | prolonged outage | X-Ray trace map |
⚠️ On NAT gateways: SAA1 lesson 3 flagged that AWS now documents "Regional NAT gateways for automatic
multi-AZ expansion", which changes the old "one per AZ" advice. I have not verified that page.
Check nat-gateways-regional.html before answering a NAT high-availability question from habit.