Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
For use during an incident. Symptom → the command that confirms it → the decision → escalation. No teaching here; each entry links to the lesson that explains it.
⚠️ Region and identifiers are yours to fill in. Every command below is read-only except where marked [WRITE].
Confirm:
aws elbv2 describe-target-health --target-group-arn <TG_ARN> \
--query 'TargetHealthDescriptions[].{Id:Target.Id,AZ:Target.AvailabilityZone,State:TargetHealth.State,Reason:TargetHealth.Reason}' \
--output table
| What you see | Decision |
|---|---|
All unhealthy, reason Target.ResponseCodeMismatch |
Health check path/port is wrong, or the app is genuinely failing. Check the app first. |
All unhealthy in one AZ |
That AZ's subnet lost its targets — its IP is pulled from DNS. Clients ignoring TTL still fail. → lesson 3 |
unused |
Target not registered, or the LB has no subnet in that AZ. |
Escalate with: target group ARN, the table above, the health check configuration, and the time the first target went unhealthy.
Near-certainly DNS caching against a withdrawn address.
Confirm:
dig +short <your-lb-dns-name> # what resolution returns now
aws elbv2 describe-load-balancers --names <LB_NAME> \
--query 'LoadBalancers[0].AvailabilityZones[].{AZ:ZoneName,Subnet:SubnetId}' --output table
Compare the count of resolved IPs against the count of enabled AZs. Fewer IPs than AZs ⇒ an AZ was pulled from DNS.
Decision: nothing to fix at the load balancer — restore healthy targets in that AZ. If this recurs and clients are known DNS-cachers, the durable fix is Global Accelerator (static anycast IPs). → lesson 3
Confirm cross-zone setting:
aws elbv2 describe-target-group-attributes --target-group-arn <TG_ARN> \
--query "Attributes[?Key=='load_balancing.cross_zone.enabled']"
aws elbv2 describe-target-health --target-group-arn <TG_ARN> \
--query 'TargetHealthDescriptions[].Target.AvailabilityZone' --output text | tr ' ' '\n' | sort | uniq -c
Decision: unbalanced target counts per AZ with cross-zone disabled. Either [WRITE] enable cross-zone load balancing, or rebalance the Auto Scaling group's subnets. → lesson 3
Confirm:
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names <ASG> \
--query 'AutoScalingGroups[0].{Min:MinSize,Max:MaxSize,Desired:DesiredCapacity,AZs:AvailabilityZones,Subnets:VPCZoneIdentifier}'
aws autoscaling describe-scaling-activities --auto-scaling-group-name <ASG> --max-items 10 \
--query 'Activities[].{Time:StartTime,Status:StatusCode,Cause:Cause}' --output table
| What you see | Decision |
|---|---|
Desired == Max |
Raise max. Working as configured. → lesson 2 |
| Activities show capacity errors | Instance type unavailable in that AZ, or a service quota. → §7 |
One AZ in AvailabilityZones |
Only one subnet attached. A subnet is zonal. → lesson 5 |
| Nothing at all | The alarm never fired — check the CloudWatch alarm, not the ASG. |
Confirm:
aws cloudwatch get-metric-statistics --namespace AWS/RDS \
--metric-name DatabaseConnections --dimensions Name=DBInstanceIdentifier,Value=<DB> \
--start-time $(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '-1 hour' +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) --period 60 --statistics Maximum --output table
aws rds describe-db-proxies --query 'DBProxies[].{Name:DBProxyName,Status:Status}' --output table
Decision: if connections track invocation count and there is no proxy, this is the Lambda connection-storm shape. Short term, reduce Lambda reserved concurrency. Durable fix: RDS Proxy (writer instance only, same VPC, not public). → lesson 5
Confirm:
aws rds describe-events --source-identifier <DB> --source-type db-instance --duration 60 \
--query 'Events[].{Time:Date,Message:Message}' --output table
Look for Multi-AZ instance failover started / completed.
Decision: failover completed but the application held dead connections. Without RDS Proxy, "a Multi-AZ failover drops every open connection" — the app must reconnect. Restart/recycle the application tier. Durable fix: RDS Proxy "preserving application connections". → lesson 5
Confirm, in the DR Region:
aws service-quotas list-service-quotas --service-code ec2 --region <DR_REGION> \
--query "Quotas[?contains(QuotaName,'On-Demand')].{Name:QuotaName,Value:Value,Adjustable:Adjustable}" \
--output table
Decision: quota ceiling, not a design fault. Request an increase now; it is not instant. "Ensure that service quotas in your DR Region are set high enough so as to not limit you from scaling up to production capacity." → lesson 5
Prevent: re-verify DR-Region quotas on a schedule, not after an incident.
Confirm:
aws xray get-trace-summaries \
--start-time $(date -u -v-15M +%s 2>/dev/null || date -u -d '-15 min' +%s) \
--end-time $(date -u +%s) \
--filter-expression 'responsetime > 3' \
--query 'TraceSummaries[].{Id:Id,Duration:Duration,Http:Http.HttpStatus}' --output table
Then open the trace map in the console for one slow trace ID.
Decision: the map names the slow downstream call. CloudWatch cannot — it reports per service, not per call across boundaries. → lesson 5
Before touching anything, answer these three:
Then, in order:
# a. DR-Region capacity headroom (see §7)
# b. replication lag — is the data current enough for the RPO?
aws rds describe-db-clusters --region <DR_REGION> \
--query 'DBClusters[].{Id:DBClusterIdentifier,Status:Status}' --output table
# c. confirm the traffic-management mechanism you will use
aws route53 list-health-checks --query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type}' --output table
Use a data plane mechanism to switch traffic. Route 53 health checks, or Application Recovery Controller routing controls. Do not change Route 53 weights or Global Accelerator traffic dials as the primary mechanism — both are control plane operations. → lesson 6
Manual trigger, automated steps.
Have these ready before opening a case:
| Always | |
|---|---|
| Account ID and Region | |
| Exact UTC timestamp the symptom started | |
| Resource ARNs (LB, target group, ASG, DB instance, proxy) | |
| The command output above, not a description of it |
| By symptom | Additional |
|---|---|
| Load balancer | target group ARN, describe-target-health output, health check config, access logs |
| Auto Scaling | ASG name, describe-scaling-activities output, launch template version |
| RDS | DB instance ID, describe-events for the window, CloudWatch DatabaseConnections + CPUUtilization |
| Quota | Region, service code, quota name, the value you need and by when |
| Latency | X-Ray trace IDs (not screenshots) and the time window |
⚠️ State the RTO/RPO you are working to. It changes what Support recommends, and it's the first thing they'll ask that you won't have thought about mid-incident.