Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Verified 2026-09-23 against the pages cited in each lesson. Items marked [unverified] were not retrieved from a page I fetched — look them up.
| What you can lose | The control |
|---|---|
| one instance | Auto Scaling group + ELB health checks |
| one AZ | subnets in ≥2 AZs; Multi-AZ RDS |
| one Region | backup/restore → pilot light → warm standby → active/active |
| one dependency | queue + DLQ + idempotent consumers |
Pick the cheapest row that meets the stated RTO/RPO. No numbers given → break the tie on operational overhead.
| Need | Service |
|---|---|
| work done once, by any worker | SQS (poll) |
| every subscriber gets every message | SNS (push, fan-out) |
| route per-event on content | EventBridge |
| ordered steps, retries, approval | Step Functions |
SNS → SQS → consumer beats SNS → consumer — the queue survives a broken consumer.
SQS: at-least-once (standard) ⇒ idempotent consumers. FIFO 3,000/s batched · 300/s unbatched, high-throughput 70,000/s. Retention 1 min–14 days, default 4 days. Max 1 MiB → claim check via S3. Long polling = no extra charge. Visibility timeout < processing time ⇒ duplicates.
SNS endpoints: SQS · Lambda · HTTP(S) · Email · Mobile push · SMS · Data Firehose · partners.
Step Functions:
| Standard | Express | |
|---|---|---|
| Semantics | exactly-once | at-least-once |
| Duration | 1 year | 5 minutes |
| Rate | 2,000/s | 100,000/s |
| Pricing | per state transition | per execution + duration |
| Patterns | Request Response, .sync, .waitForTaskToken | Request Response only |
⚠️ Human approval / wait-for-callback ⇒ Standard.
Vertical = performance. Horizontal = availability. Horizontal requires statelessness.
| State | Goes to |
|---|---|
| session | ElastiCache / DynamoDB / signed cookie |
| uploads | S3 |
| shared files | EFS |
| data | RDS / Aurora / DynamoDB |
Sticky sessions = workaround. Right answer only when "application changes are not possible".
Auto Scaling group = min / max / desired. Scaling policies move desired.
⚠️ "Auto Scaling is a control plane activity" → don't depend on it in a failover. Statically stable alternative = hot standby. [unverified] scaling policy types (target tracking / step / simple / scheduled / predictive)
| ALB | NLB | GWLB | |
|---|---|---|---|
| Layer | 7 | 4 | 3 (per exam guide) |
| Protocols | HTTP/S | TCP, UDP, TCP_UDP, TLS, QUIC | GENEVE |
| Routes on | content | flow hash | — |
| Static IP | ✗ | ✅ 1 EIP per subnet | — |
| WAF attaches | ✅ | ✗ | ✗ |
ALB only: path conditions · host conditions · HTTP header / method / query / source IP · redirect · fixed response · Lambda targets · user authentication. Default algorithm round robin; alternative least outstanding requests.
NLB only: static IPs · "millions of requests per second" · targets by IP, including outside the VPC · QUIC.
⚠️ Cross-zone OFF by default at the node level: "each load balancer node distributes traffic across the registered targets in its Availability Zone only." 2 targets in AZ-A + 8 in AZ-B ⇒ the 2 get 25% each. Fix: enable cross-zone or balance target counts. ⚠️ No healthy target in an AZ ⇒ its IP is pulled from DNS; clients ignoring TTL still fail. [unverified] per-type cross-zone defaults and charging.
Global Accelerator — static anycast IPs, traffic across "multiple load balancers in one or more AWS Regions", "avoids caching issues that can occur with DNS systems". → the answer to "failover is slow because clients cache DNS".
ORCHESTRATOR: ECS | EKS ← ecosystem
CAPACITY: Fargate | EC2 ← control vs overhead
"ECS vs Fargate" is a category error.
ECS: task definition = blueprint · cluster = infrastructure · task = runs and stops · service = long-running, desired count. Cluster Auto Scaling (hosts) vs Service Auto Scaling (tasks). Capacity: Fargate · EC2 · ECS Managed Instances · ECS Anywhere (on-prem).
Fargate: removes "choose server types, decide when to scale your clusters, or optimize cluster
packing". Own isolation boundary per task. Fargate Spot = 2-minute warning (verified).
⚠️ Target group type must be ip, not instance. Platform patching retires running tasks.
Doesn't support all task-definition parameters → GPU/privileged ⇒ EC2.
EKS: "certified Kubernetes-conformant… without refactoring" — that's the only durable reason to pick it. EKS standard = AWS runs control plane; EKS Auto Mode = AWS runs nodes too. Per-cluster pricing (ECS has none for the orchestrator). EKS Anywhere / Hybrid Nodes for on-prem.
Lambda: 15 min max · one request per execution environment · no guaranteed state across invocations · 200+ triggers · per-request + GB-seconds. Lambda MicroVMs = 8 h, own binaries, state across suspend — not in the SAA-C03 guide.
| Highly available | Fault tolerant | |
|---|---|---|
| recovers fast (brief gap) | no interruption | |
| Example | Multi-AZ RDS | multiple active instances behind an ELB |
⚠️ Multi-AZ standby serves NO reads. Reads ⇒ read replicas (async, manual promotion). Both problems ⇒ both features.
Route 53 failover:
RDS Proxy — pools connections; "automatically connecting to a standby DB instance while preserving application connections"; "queues or throttles" then "sheds load". Lambda + RDS connection storm ⇒ RDS Proxy. 20 proxies/account · 200 secrets/proxy · writer only · same VPC · not public.
Service quotas — "Ensure that service quotas in your DR Region are set high enough". An invisible SPOF.
X-Ray — trace map across service boundaries → "bottlenecks, latency spikes". CloudWatch says a service is slow; X-Ray says which call.
[unverified] regional NAT gateways / multi-AZ expansion — check before answering NAT HA from habit.
RPO = data loss (set by replication/backup frequency). RTO = downtime (set by what must be built).
| Strategy | DR Region state | Can it serve now? |
|---|---|---|
| Backup & restore | data only; redeploy everything | no |
| Pilot light | data replicating, core infra on, servers off | no |
| Warm standby | everything on, scaled down | yes, slowly |
| Multi-site active/active | live, serving | "no such thing as failover" |
| hot standby | full capacity, active/passive | yes — statically stable |
⚠️ Pilot light vs warm standby: "pilot light cannot process requests without additional action taken first, whereas warm standby can handle traffic (at reduced capacity levels) immediately."
Replication services: S3 Replication · RDS read replicas · Aurora global database (sub-second; promote < 1 minute) · DynamoDB global tables (last writer wins) · DocumentDB global clusters · ElastiCache Global Datastore.
⚠️ Replication ≠ backup. "may not protect against… data corruption or malicious attack."
Active/active writes: write global (Aurora, write forwarding) · write local (DynamoDB global tables) · write partitioned (S3 bi-directional + replica modification sync). Reads = read local.
"For maximum resiliency, you should use only data plane operations as part of your failover operation. … data planes typically have higher availability design goals than the control planes."
| Mechanism | Plane |
|---|---|
| Route 53 health checks → DNS failover | data ✅ |
| ARC routing controls (health checks as on/off switches) | data ✅ |
| Route 53 weight changes | control ⚠️ |
| Global Accelerator traffic dials | control ⚠️ |
| Auto Scaling to reach capacity | control ⚠️ |
| AWS Backup restore | control ⚠️ |
Failover trigger: "Automatically initiated failover … should be used with caution" — false alarms cost you. → manual trigger, automated steps, "like the push of a button".
CloudFront origin failover is per request, not a sticky switch-over.
Testing: "critical to regularly assess and test" → AWS Resilience Hub. For active/active ask: "Can the other Region(s) handle all the traffic?" Two Regions at 60% cannot absorb each other.