Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Domain 2 is 26% of scored content — the second-largest domain. And it is the one where the exam's "choose the best answer" framing hurts most, because resilience is a spectrum with a price tag. Four options can all be more resilient than doing nothing, and only one matches the recovery objective and budget the stem actually stated.
The domain has only two task statements, and they are doing different jobs:
Most people study 2.2 and skim 2.1. That's backwards — 2.1 has sixteen "Knowledge of" items to 2.2's twelve, and decoupling questions are more common than DR questions.
Resilience is bought in units of "what can I lose?" — and every unit has a price.
WHAT YOU CAN LOSE WHAT IT COSTS THE CONTROL
──────────────── ──────────── ───────────
one instance → another instance → Auto Scaling + ELB health checks
one Availability Zone → capacity in 2+ AZs → multi-AZ subnets, Multi-AZ RDS
one Region → a second footprint → backup/restore → pilot light →
warm standby → active/active
one dependency → a queue and a retry → SQS, DLQ, idempotent consumers
The exam's job is to make you pick the cheapest row that meets the stated requirement. Reaching for multi-Region active/active when the stem says "development environment" is as wrong as reaching for a single instance when it says "must survive a Regional outage".
⚠️ The tell is always RTO and RPO. If the stem gives you numbers, they select the row. If it gives you "minimal downtime" with no number, the tie-break is operational overhead and cost — Domain 1's rule, applied here.
| # | Lesson | Task statement | Read | Listen |
|---|---|---|---|---|
| 1 | Decoupling — queues, topics, events and workflows | 2.1 | 30 min | 11 min |
| 2 | Scaling — horizontal, vertical, and why state is the constraint | 2.1 | 26 min | 9 min |
| 3 | Load balancing and multi-tier architecture | 2.1 | 26 min | 11 min |
| 4 | Containers and serverless — ECS, EKS, Fargate, Lambda | 2.1 | 28 min | 12 min |
| 5 | High availability — surviving an instance, an AZ, a dependency | 2.2 | 30 min | 11 min |
| 6 | Disaster recovery — the four strategies and the data plane rule | 2.2 | 32 min | 13 min |
Then: Cheat sheet · Lab · Quiz · Interview · Runbook
From the SAA-C03 exam guide — Content Domain 2 (verified 2026-09-22), Design Resilient Architectures — 26% of scored content:
| Task statement | This module |
|---|---|
| 2.1 Design scalable and loosely coupled architectures | lessons 1–4 |
| 2.2 Design highly available and/or fault-tolerant architectures | lessons 5–6 |
This module cites the HTML exam guide at
docs.aws.amazon.com/aws-certification/latest/solutions-architect-associate-03/, not the
d1.awsstatic.com PDF. The HTML is maintained and is demonstrably ahead of that PDF — see
../PLAN.md for the diff.
Two places where a page I fetched did not contain something examinable are flagged inline with ⚠️ and the page to read. The largest is the ALB/NLB/Gateway Load Balancer feature comparison: the Elastic Load Balancing "What is" page links to per-type user guides rather than carrying a comparison table, so lesson 3 builds the comparison from the exam guide's own wording plus the per-type guides, and says so.
Facts verified 2026-09-22 against the pages cited in each lesson.
Keeps playing into the following modules — 208 min from here to the end of certification prep.