AWS Training
Modules Listen All tracks

← Design Resilient Architectures

SAA2 Lab — break it, watch it heal

You will need: AWS CLI v2, jq, a default VPC with at least two subnets in different AZs (or your own), and permission to create EC2, ELB, Auto Scaling, SQS and SNS resources.

Write your answers down. Questions are numbered Q1…Q20.

⚠️ Safety: everything here is cheap and reversible. The one thing to watch is the ALB, which bills hourly whether or not you use it. Run the teardown today.

export LAB=saa2-$(date +%Y%m%d)-$RANDOM
export REGION=$(aws configure get region)
export VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true \
  --query 'Vpcs[0].VpcId' --output text)
# two subnets in DIFFERENT AZs
read SUB_A AZ_A SUB_B AZ_B <<<$(aws ec2 describe-subnets --filters Name=vpc-id,Values=$VPC \
  --query 'Subnets | sort_by(@,&AvailabilityZone) | [0].[SubnetId,AvailabilityZone] | join(` `,@)' --output text)" "$(aws ec2 describe-subnets --filters Name=vpc-id,Values=$VPC \
  --query 'Subnets | sort_by(@,&AvailabilityZone) | [-1].[SubnetId,AvailabilityZone] | join(` `,@)' --output text)
echo "LAB=$LAB VPC=$VPC"
echo "A: $SUB_A ($AZ_A)   B: $SUB_B ($AZ_B)"

Q1. Are AZ_A and AZ_B different? If not, find two subnets in different AZs before continuing. Then state, in one sentence, why this lab cannot proceed with one subnet — quote the VPC FAQ.


Part 1 — Self-healing without scaling (25 min)

Lesson 2 claims an Auto Scaling group with min = max = desired replaces failures. Prove it.

AMI=$(aws ssm get-parameters --names \
  /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
  --query 'Parameters[0].Value' --output text)

aws ec2 create-launch-template --launch-template-name $LAB-lt \
  --launch-template-data "{\"ImageId\":\"$AMI\",\"InstanceType\":\"t3.micro\"}" >/dev/null

aws autoscaling create-auto-scaling-group --auto-scaling-group-name $LAB-asg \
  --launch-template LaunchTemplateName=$LAB-lt,Version='$Latest' \
  --min-size 2 --max-size 2 --desired-capacity 2 \
  --vpc-zone-identifier "$SUB_A,$SUB_B"

sleep 60
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names $LAB-asg \
  --query 'AutoScalingGroups[0].Instances[].{Id:InstanceId,AZ:AvailabilityZone,State:LifecycleState,Health:HealthStatus}' \
  --output table

Q2. Record the two instances and their AZs. Did Auto Scaling place one in each? Quote the sentence from the Auto Scaling documentation that predicted this.

Now kill one:

VICTIM=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names $LAB-asg \
  --query 'AutoScalingGroups[0].Instances[0].InstanceId' --output text)
echo "terminating $VICTIM"
aws ec2 terminate-instances --instance-ids $VICTIM >/dev/null

Watch (re-run every 30 seconds until it settles):

aws autoscaling describe-scaling-activities --auto-scaling-group-name $LAB-asg --max-items 5 \
  --query 'Activities[].{Time:StartTime,Status:StatusCode,Cause:Cause}' --output table

Q3. Paste the Cause text for the replacement activity. In AWS's own words, what triggered it?

Q4. This group has no scaling policy at all. Write one sentence explaining what it is for, and name the requirement phrasing in an exam stem that this answers.

Q5. Now the max trap. Set desired to 4 and observe:

aws autoscaling set-desired-capacity --auto-scaling-group-name $LAB-asg --desired-capacity 4 2>&1 | tail -3

Paste the error. Which of min/max/desired stopped you, and what is the one-line fix?


Part 2 — Cross-zone load balancing, the hard way (35 min)

This is the deliberate failure: you will cause uneven distribution and then diagnose it.

SG=$(aws ec2 create-security-group --group-name $LAB-sg --description "$LAB" \
  --vpc-id $VPC --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id $SG \
  --protocol tcp --port 80 --cidr 0.0.0.0/0 >/dev/null

ALB=$(aws elbv2 create-load-balancer --name $LAB-alb --type application \
  --subnets $SUB_A $SUB_B --security-groups $SG \
  --query 'LoadBalancers[0].LoadBalancerArn' --output text)

TG=$(aws elbv2 create-target-group --name $LAB-tg --protocol HTTP --port 80 \
  --vpc-id $VPC --target-type instance --query 'TargetGroups[0].TargetGroupArn' --output text)

aws elbv2 create-listener --load-balancer-arn $ALB --protocol HTTP --port 80 \
  --default-actions Type=forward,TargetGroupArn=$TG >/dev/null

Q6. Check the cross-zone setting on the target group, and on the load balancer:

aws elbv2 describe-target-group-attributes --target-group-arn $TG \
  --query "Attributes[?contains(Key,'cross_zone')]"
aws elbv2 describe-load-balancer-attributes --load-balancer-arn $ALB \
  --query "Attributes[?contains(Key,'cross_zone')]"

Record both. Lesson 3 flagged the per-type default as unverified — you have now verified it for your own account and load balancer type. Write down what you found, and what it means for the uneven-distribution scenario.

Q7. Attach the ASG to the target group and deliberately unbalance the AZs:

aws autoscaling attach-load-balancer-target-groups \
  --auto-scaling-group-name $LAB-asg --target-group-arns $TG
# force everything into ONE AZ
aws autoscaling update-auto-scaling-group --auto-scaling-group-name $LAB-asg \
  --vpc-zone-identifier "$SUB_A" --min-size 2 --max-size 2 --desired-capacity 2
sleep 90
aws elbv2 describe-target-health --target-group-arn $TG \
  --query 'TargetHealthDescriptions[].{Id:Target.Id,AZ:Target.AvailabilityZone,State:TargetHealth.State,Reason:TargetHealth.Reason}' \
  --output table

Record the table. All targets are now in one AZ, but the load balancer still has a subnet in two.

Q8. Resolve the ALB's DNS name and count the addresses:

DNS=$(aws elbv2 describe-load-balancers --load-balancer-arns $ALB \
  --query 'LoadBalancers[0].DNSName' --output text)
dig +short $DNS

How many IP addresses came back, and how many AZs does the load balancer have enabled? Now read the NLB introduction quote from lesson 3 about DNS removal. Does what you observed match it? If it doesn't, say so — that is a legitimate finding, and note which load balancer type the quote was about.

Q9. Targets are unhealthy (nothing is listening on port 80 — we never installed a web server). Explain why that is useful for this lab, and what the Reason field told you.

Q10. Given 2 targets in AZ-A and 8 in AZ-B with cross-zone disabled, compute the share of traffic each individual target receives. Show the arithmetic.

Q11. Name the two legitimate fixes for uneven distribution, and quote the AWS sentence that recommends the second one.


Part 3 — Decoupling, and the duplicate you cause (30 min)

QURL=$(aws sqs create-queue --queue-name $LAB-q --query QueueUrl --output text)
DLQ=$(aws sqs create-queue --queue-name $LAB-dlq --query QueueUrl --output text)
DLQ_ARN=$(aws sqs get-queue-attributes --queue-url $DLQ \
  --attribute-names QueueArn --query 'Attributes.QueueArn' --output text)

Q12. Wire up the dead-letter queue with a redrive policy of 2 receives, then confirm it:

aws sqs set-queue-attributes --queue-url $QURL --attributes \
  "{\"RedrivePolicy\":\"{\\\"deadLetterTargetArn\\\":\\\"$DLQ_ARN\\\",\\\"maxReceiveCount\\\":\\\"2\\\"}\",\"VisibilityTimeout\":\"5\"}"
aws sqs get-queue-attributes --queue-url $QURL \
  --attribute-names RedrivePolicy VisibilityTimeout --query Attributes

Record the output. What does maxReceiveCount of 2 mean in plain words?

Q13. Now be the poison message. Send one, receive it twice without deleting it, and find it in the DLQ:

aws sqs send-message --queue-url $QURL --message-body "poison" >/dev/null
for i in 1 2 3; do
  echo "--- receive $i ---"
  aws sqs receive-message --queue-url $QURL --wait-time-seconds 10 \
    --query 'Messages[0].MessageId' --output text
  sleep 6   # let the 5s visibility timeout expire
done
aws sqs receive-message --queue-url $DLQ --wait-time-seconds 10 \
  --query 'Messages[0].Body' --output text

On which receive did the message stop coming back, and where did it end up? Explain what would have happened without a dead-letter queue, including how long it would have gone on.

Q14. You just reproduced duplicate delivery on purpose. State the design rule that follows, and the two configuration values that were in tension.

Q15. Fan-out. Create a topic, subscribe two queues, publish once:

TOPIC=$(aws sns create-topic --name $LAB-topic --query TopicArn --output text)
for N in a b; do
  Q=$(aws sqs create-queue --queue-name $LAB-fan-$N --query QueueUrl --output text)
  QA=$(aws sqs get-queue-attributes --queue-url $Q --attribute-names QueueArn \
        --query 'Attributes.QueueArn' --output text)
  aws sqs set-queue-attributes --queue-url $Q --attributes \
    "{\"Policy\":\"{\\\"Version\\\":\\\"2012-10-17\\\",\\\"Statement\\\":[{\\\"Effect\\\":\\\"Allow\\\",\\\"Principal\\\":{\\\"Service\\\":\\\"sns.amazonaws.com\\\"},\\\"Action\\\":\\\"sqs:SendMessage\\\",\\\"Resource\\\":\\\"$QA\\\",\\\"Condition\\\":{\\\"ArnEquals\\\":{\\\"aws:SourceArn\\\":\\\"$TOPIC\\\"}}}]}\"}"
  aws sns subscribe --topic-arn $TOPIC --protocol sqs --notification-endpoint $QA >/dev/null
  eval "Q_$N=$Q"
done

aws sns publish --topic-arn $TOPIC --message "one event" >/dev/null
sleep 5
for N in a b; do
  eval "Q=\$Q_$N"
  echo -n "queue $N: "
  aws sqs receive-message --queue-url $Q --wait-time-seconds 5 \
    --query 'length(Messages)' --output text
done

Q16. How many copies of one published message exist? Now answer the design question: if queue b's consumer were down for an hour, would the message survive? Would it have survived with a direct SNS→Lambda subscription instead? Explain the difference in one sentence.

Q17. You had to attach a queue policy allowing sns.amazonaws.com to sqs:SendMessage. Which lesson from Domain 1 does that demonstrate, and what would have happened without it?


Part 4 — Reading the plan, not the console (20 min, no writes)

No commands. These are the questions an interviewer or an exam stem will actually ask.

Q18. For each requirement, name the DR strategy and justify it in one sentence with a quote:

# Requirement
a RPO 24 h, RTO 24 h, minimum cost
b RPO 5 min, RTO 30 min, cost-conscious
c RPO seconds, RTO minutes, cannot wait for deployment
d RTO and RPO near zero, cost secondary

Q19. Your DR plan uses Auto Scaling to reach production capacity in the recovery Region, and shifts traffic by changing Route 53 weights. Identify both control-plane dependencies, quote AWS on each, and give the data-plane alternative.

Q20. An engineer runs a destructive DELETE against production at 09:00. You replicate continuously to a second Region with sub-second lag. State the time the problem reaches your DR Region, and name the three controls that would have helped (one of them is from SAA1).


Teardown — today

aws elbv2 delete-listener --listener-arn $(aws elbv2 describe-listeners \
  --load-balancer-arn $ALB --query 'Listeners[0].ListenerArn' --output text) 2>/dev/null
aws elbv2 delete-load-balancer --load-balancer-arn $ALB
sleep 30
aws elbv2 delete-target-group --target-group-arn $TG

aws autoscaling update-auto-scaling-group --auto-scaling-group-name $LAB-asg \
  --min-size 0 --max-size 0 --desired-capacity 0
aws autoscaling delete-auto-scaling-group --auto-scaling-group-name $LAB-asg --force-delete
aws ec2 delete-launch-template --launch-template-name $LAB-lt >/dev/null

for U in $QURL $DLQ $Q_a $Q_b; do aws sqs delete-queue --queue-url $U 2>/dev/null; done
aws sns delete-topic --topic-arn $TOPIC

sleep 60
aws ec2 delete-security-group --group-id $SG

Verify it — don't assume it:

aws elbv2 describe-load-balancers --query "LoadBalancers[?starts_with(LoadBalancerName,'$LAB')].LoadBalancerName"
aws autoscaling describe-auto-scaling-groups --query "AutoScalingGroups[?starts_with(AutoScalingGroupName,'$LAB')].AutoScalingGroupName"
aws sqs list-queues --queue-name-prefix $LAB
aws ec2 describe-instances --filters Name=tag:aws:autoscaling:groupName,Values=$LAB-asg \
  Name=instance-state-name,Values=running --query 'Reservations[].Instances[].InstanceId'

All four should be empty. If the security group deletion failed, wait a few minutes — ENIs take time to detach — and retry.


Done when you can

Facilitator notes