AWS Training
Modules Listen All tracks

← Design Resilient Architectures

SAA2 Runbook — resilience incidents

For use during an incident. Symptom → the command that confirms it → the decision → escalation. No teaching here; each entry links to the lesson that explains it.

⚠️ Region and identifiers are yours to fill in. Every command below is read-only except where marked [WRITE].


1. "The site is down but the instances look fine"

Confirm:

aws elbv2 describe-target-health --target-group-arn <TG_ARN> \
  --query 'TargetHealthDescriptions[].{Id:Target.Id,AZ:Target.AvailabilityZone,State:TargetHealth.State,Reason:TargetHealth.Reason}' \
  --output table
What you see Decision
All unhealthy, reason Target.ResponseCodeMismatch Health check path/port is wrong, or the app is genuinely failing. Check the app first.
All unhealthy in one AZ That AZ's subnet lost its targets — its IP is pulled from DNS. Clients ignoring TTL still fail. → lesson 3
unused Target not registered, or the LB has no subnet in that AZ.

Escalate with: target group ARN, the table above, the health check configuration, and the time the first target went unhealthy.


2. "Some clients get errors, most are fine"

Near-certainly DNS caching against a withdrawn address.

Confirm:

dig +short <your-lb-dns-name>          # what resolution returns now
aws elbv2 describe-load-balancers --names <LB_NAME> \
  --query 'LoadBalancers[0].AvailabilityZones[].{AZ:ZoneName,Subnet:SubnetId}' --output table

Compare the count of resolved IPs against the count of enabled AZs. Fewer IPs than AZs ⇒ an AZ was pulled from DNS.

Decision: nothing to fix at the load balancer — restore healthy targets in that AZ. If this recurs and clients are known DNS-cachers, the durable fix is Global Accelerator (static anycast IPs). → lesson 3


3. "Two instances are melting, eight are idle"

Confirm cross-zone setting:

aws elbv2 describe-target-group-attributes --target-group-arn <TG_ARN> \
  --query "Attributes[?Key=='load_balancing.cross_zone.enabled']"
aws elbv2 describe-target-health --target-group-arn <TG_ARN> \
  --query 'TargetHealthDescriptions[].Target.AvailabilityZone' --output text | tr ' ' '\n' | sort | uniq -c

Decision: unbalanced target counts per AZ with cross-zone disabled. Either [WRITE] enable cross-zone load balancing, or rebalance the Auto Scaling group's subnets. → lesson 3


4. "Auto Scaling isn't scaling"

Confirm:

aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names <ASG> \
  --query 'AutoScalingGroups[0].{Min:MinSize,Max:MaxSize,Desired:DesiredCapacity,AZs:AvailabilityZones,Subnets:VPCZoneIdentifier}'
aws autoscaling describe-scaling-activities --auto-scaling-group-name <ASG> --max-items 10 \
  --query 'Activities[].{Time:StartTime,Status:StatusCode,Cause:Cause}' --output table
What you see Decision
Desired == Max Raise max. Working as configured. → lesson 2
Activities show capacity errors Instance type unavailable in that AZ, or a service quota. → §7
One AZ in AvailabilityZones Only one subnet attached. A subnet is zonal. → lesson 5
Nothing at all The alarm never fired — check the CloudWatch alarm, not the ASG.

5. "Database connections exhausted"

Confirm:

aws cloudwatch get-metric-statistics --namespace AWS/RDS \
  --metric-name DatabaseConnections --dimensions Name=DBInstanceIdentifier,Value=<DB> \
  --start-time $(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '-1 hour' +%Y-%m-%dT%H:%M:%SZ) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) --period 60 --statistics Maximum --output table
aws rds describe-db-proxies --query 'DBProxies[].{Name:DBProxyName,Status:Status}' --output table

Decision: if connections track invocation count and there is no proxy, this is the Lambda connection-storm shape. Short term, reduce Lambda reserved concurrency. Durable fix: RDS Proxy (writer instance only, same VPC, not public). → lesson 5


6. "Failover happened but the app didn't recover"

Confirm:

aws rds describe-events --source-identifier <DB> --source-type db-instance --duration 60 \
  --query 'Events[].{Time:Date,Message:Message}' --output table

Look for Multi-AZ instance failover started / completed.

Decision: failover completed but the application held dead connections. Without RDS Proxy, "a Multi-AZ failover drops every open connection" — the app must reconnect. Restart/recycle the application tier. Durable fix: RDS Proxy "preserving application connections". → lesson 5


7. "DR failover ran but we only got a fraction of the capacity"

Confirm, in the DR Region:

aws service-quotas list-service-quotas --service-code ec2 --region <DR_REGION> \
  --query "Quotas[?contains(QuotaName,'On-Demand')].{Name:QuotaName,Value:Value,Adjustable:Adjustable}" \
  --output table

Decision: quota ceiling, not a design fault. Request an increase now; it is not instant. "Ensure that service quotas in your DR Region are set high enough so as to not limit you from scaling up to production capacity." → lesson 5

Prevent: re-verify DR-Region quotas on a schedule, not after an incident.


8. "Something in the microservice chain is slow"

Confirm:

aws xray get-trace-summaries \
  --start-time $(date -u -v-15M +%s 2>/dev/null || date -u -d '-15 min' +%s) \
  --end-time $(date -u +%s) \
  --filter-expression 'responsetime > 3' \
  --query 'TraceSummaries[].{Id:Id,Duration:Duration,Http:Http.HttpStatus}' --output table

Then open the trace map in the console for one slow trace ID.

Decision: the map names the slow downstream call. CloudWatch cannot — it reports per service, not per call across boundaries. → lesson 5


9. "We need to fail over to the DR Region"

Before touching anything, answer these three:

  1. Is this a Region event or a data event? Replication copies corruption. If data was deleted or corrupted, failing over replicates the problem — restore from a point-in-time backup instead. → lesson 6
  2. What is our strategy? Pilot light (servers must be switched on first) or warm standby (scale up only)? The test: can the DR Region serve a request right now?
  3. Is the trigger a false alarm? "If you fail over when you don't need to (false alarm), then you incur those losses."

Then, in order:

# a. DR-Region capacity headroom (see §7)
# b. replication lag — is the data current enough for the RPO?
aws rds describe-db-clusters --region <DR_REGION> \
  --query 'DBClusters[].{Id:DBClusterIdentifier,Status:Status}' --output table
# c. confirm the traffic-management mechanism you will use
aws route53 list-health-checks --query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type}' --output table

Use a data plane mechanism to switch traffic. Route 53 health checks, or Application Recovery Controller routing controls. Do not change Route 53 weights or Global Accelerator traffic dials as the primary mechanism — both are control plane operations. → lesson 6

Manual trigger, automated steps.


Escalation kit — what AWS Support will ask for

Have these ready before opening a case:

Always
Account ID and Region
Exact UTC timestamp the symptom started
Resource ARNs (LB, target group, ASG, DB instance, proxy)
The command output above, not a description of it
By symptom Additional
Load balancer target group ARN, describe-target-health output, health check config, access logs
Auto Scaling ASG name, describe-scaling-activities output, launch template version
RDS DB instance ID, describe-events for the window, CloudWatch DatabaseConnections + CPUUtilization
Quota Region, service code, quota name, the value you need and by when
Latency X-Ray trace IDs (not screenshots) and the time window

⚠️ State the RTO/RPO you are working to. It changes what Support recommends, and it's the first thing they'll ask that you won't have thought about mid-incident.