AWS Training
Modules Listen All tracks

← Design Resilient Architectures

SAA2 Cheat sheet — Design Resilient Architectures (Domain 2, 26%)

Verified 2026-09-23 against the pages cited in each lesson. Items marked [unverified] were not retrieved from a page I fetched — look them up.

The buying ladder

What you can lose The control
one instance Auto Scaling group + ELB health checks
one AZ subnets in ≥2 AZs; Multi-AZ RDS
one Region backup/restore → pilot light → warm standby → active/active
one dependency queue + DLQ + idempotent consumers

Pick the cheapest row that meets the stated RTO/RPO. No numbers given → break the tie on operational overhead.

Decoupling

Need Service
work done once, by any worker SQS (poll)
every subscriber gets every message SNS (push, fan-out)
route per-event on content EventBridge
ordered steps, retries, approval Step Functions

SNS → SQS → consumer beats SNS → consumer — the queue survives a broken consumer.

SQS: at-least-once (standard) ⇒ idempotent consumers. FIFO 3,000/s batched · 300/s unbatched, high-throughput 70,000/s. Retention 1 min–14 days, default 4 days. Max 1 MiB → claim check via S3. Long polling = no extra charge. Visibility timeout < processing time ⇒ duplicates.

SNS endpoints: SQS · Lambda · HTTP(S) · Email · Mobile push · SMS · Data Firehose · partners.

Step Functions:

Standard Express
Semantics exactly-once at-least-once
Duration 1 year 5 minutes
Rate 2,000/s 100,000/s
Pricing per state transition per execution + duration
Patterns Request Response, .sync, .waitForTaskToken Request Response only

⚠️ Human approval / wait-for-callback ⇒ Standard.

Scaling

Vertical = performance. Horizontal = availability. Horizontal requires statelessness.

State Goes to
session ElastiCache / DynamoDB / signed cookie
uploads S3
shared files EFS
data RDS / Aurora / DynamoDB

Sticky sessions = workaround. Right answer only when "application changes are not possible".

Auto Scaling group = min / max / desired. Scaling policies move desired.

⚠️ "Auto Scaling is a control plane activity" → don't depend on it in a failover. Statically stable alternative = hot standby. [unverified] scaling policy types (target tracking / step / simple / scheduled / predictive)

Load balancing — it's an OSI layer question

ALB NLB GWLB
Layer 7 4 3 (per exam guide)
Protocols HTTP/S TCP, UDP, TCP_UDP, TLS, QUIC GENEVE
Routes on content flow hash —
Static IP ✗ ✅ 1 EIP per subnet —
WAF attaches ✅ ✗ ✗

ALB only: path conditions · host conditions · HTTP header / method / query / source IP · redirect · fixed response · Lambda targets · user authentication. Default algorithm round robin; alternative least outstanding requests.

NLB only: static IPs · "millions of requests per second" · targets by IP, including outside the VPC · QUIC.

⚠️ Cross-zone OFF by default at the node level: "each load balancer node distributes traffic across the registered targets in its Availability Zone only." 2 targets in AZ-A + 8 in AZ-B ⇒ the 2 get 25% each. Fix: enable cross-zone or balance target counts. ⚠️ No healthy target in an AZ ⇒ its IP is pulled from DNS; clients ignoring TTL still fail. [unverified] per-type cross-zone defaults and charging.

Global Accelerator — static anycast IPs, traffic across "multiple load balancers in one or more AWS Regions", "avoids caching issues that can occur with DNS systems". → the answer to "failover is slow because clients cache DNS".

Containers & serverless — two axes

ORCHESTRATOR:  ECS  |  EKS          ← ecosystem
CAPACITY:      Fargate  |  EC2      ← control vs overhead

"ECS vs Fargate" is a category error.

ECS: task definition = blueprint · cluster = infrastructure · task = runs and stops · service = long-running, desired count. Cluster Auto Scaling (hosts) vs Service Auto Scaling (tasks). Capacity: Fargate · EC2 · ECS Managed Instances · ECS Anywhere (on-prem).

Fargate: removes "choose server types, decide when to scale your clusters, or optimize cluster packing". Own isolation boundary per task. Fargate Spot = 2-minute warning (verified). ⚠️ Target group type must be ip, not instance. Platform patching retires running tasks. Doesn't support all task-definition parameters → GPU/privileged ⇒ EC2.

EKS: "certified Kubernetes-conformant… without refactoring" — that's the only durable reason to pick it. EKS standard = AWS runs control plane; EKS Auto Mode = AWS runs nodes too. Per-cluster pricing (ECS has none for the orchestrator). EKS Anywhere / Hybrid Nodes for on-prem.

Lambda: 15 min max · one request per execution environment · no guaranteed state across invocations · 200+ triggers · per-request + GB-seconds. Lambda MicroVMs = 8 h, own binaries, state across suspend — not in the SAA-C03 guide.

High availability

Highly available Fault tolerant
recovers fast (brief gap) no interruption
Example Multi-AZ RDS multiple active instances behind an ELB

⚠️ Multi-AZ standby serves NO reads. Reads ⇒ read replicas (async, manual promotion). Both problems ⇒ both features.

Route 53 failover:

RDS Proxy — pools connections; "automatically connecting to a standby DB instance while preserving application connections"; "queues or throttles" then "sheds load". Lambda + RDS connection storm ⇒ RDS Proxy. 20 proxies/account · 200 secrets/proxy · writer only · same VPC · not public.

Service quotas — "Ensure that service quotas in your DR Region are set high enough". An invisible SPOF.

X-Ray — trace map across service boundaries → "bottlenecks, latency spikes". CloudWatch says a service is slow; X-Ray says which call.

[unverified] regional NAT gateways / multi-AZ expansion — check before answering NAT HA from habit.

Disaster recovery

RPO = data loss (set by replication/backup frequency). RTO = downtime (set by what must be built).

Strategy DR Region state Can it serve now?
Backup & restore data only; redeploy everything no
Pilot light data replicating, core infra on, servers off no
Warm standby everything on, scaled down yes, slowly
Multi-site active/active live, serving "no such thing as failover"
hot standby full capacity, active/passive yes — statically stable

⚠️ Pilot light vs warm standby: "pilot light cannot process requests without additional action taken first, whereas warm standby can handle traffic (at reduced capacity levels) immediately."

Replication services: S3 Replication · RDS read replicas · Aurora global database (sub-second; promote < 1 minute) · DynamoDB global tables (last writer wins) · DocumentDB global clusters · ElastiCache Global Datastore.

⚠️ Replication ≠ backup. "may not protect against… data corruption or malicious attack."

Active/active writes: write global (Aurora, write forwarding) · write local (DynamoDB global tables) · write partitioned (S3 bi-directional + replica modification sync). Reads = read local.

⭐ The data plane rule

"For maximum resiliency, you should use only data plane operations as part of your failover operation. … data planes typically have higher availability design goals than the control planes."

Mechanism Plane
Route 53 health checks → DNS failover data ✅
ARC routing controls (health checks as on/off switches) data ✅
Route 53 weight changes control ⚠️
Global Accelerator traffic dials control ⚠️
Auto Scaling to reach capacity control ⚠️
AWS Backup restore control ⚠️

Failover trigger: "Automatically initiated failover … should be used with caution" — false alarms cost you. → manual trigger, automated steps, "like the push of a button".

CloudFront origin failover is per request, not a sticky switch-over.

Testing: "critical to regularly assess and test" → AWS Resilience Hub. For active/active ask: "Can the other Region(s) handle all the traffic?" Two Regions at 60% cannot absorb each other.

Four things to carry in

  1. Resilience is bought in units of what you can lose — pick the cheapest row that meets the stated RTO/RPO.
  2. A queue converts rejection into delay. Delay is the cheaper failure.
  3. Multi-AZ is availability; read replicas are scaling. Different features.
  4. Replication is not backup — it copies your mistakes perfectly.