Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
You will need: AWS CLI v2, jq, a default VPC with at least two subnets in different AZs (or
your own), and permission to create EC2, ELB, Auto Scaling, SQS and SNS resources.
Write your answers down. Questions are numbered Q1…Q20.
⚠️ Safety: everything here is cheap and reversible. The one thing to watch is the ALB, which bills hourly whether or not you use it. Run the teardown today.
export LAB=saa2-$(date +%Y%m%d)-$RANDOM
export REGION=$(aws configure get region)
export VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true \
--query 'Vpcs[0].VpcId' --output text)
# two subnets in DIFFERENT AZs
read SUB_A AZ_A SUB_B AZ_B <<<$(aws ec2 describe-subnets --filters Name=vpc-id,Values=$VPC \
--query 'Subnets | sort_by(@,&AvailabilityZone) | [0].[SubnetId,AvailabilityZone] | join(` `,@)' --output text)" "$(aws ec2 describe-subnets --filters Name=vpc-id,Values=$VPC \
--query 'Subnets | sort_by(@,&AvailabilityZone) | [-1].[SubnetId,AvailabilityZone] | join(` `,@)' --output text)
echo "LAB=$LAB VPC=$VPC"
echo "A: $SUB_A ($AZ_A) B: $SUB_B ($AZ_B)"
Q1. Are AZ_A and AZ_B different? If not, find two subnets in different AZs before continuing.
Then state, in one sentence, why this lab cannot proceed with one subnet — quote the VPC FAQ.
Lesson 2 claims an Auto Scaling group with min = max = desired replaces failures. Prove it.
AMI=$(aws ssm get-parameters --names \
/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
--query 'Parameters[0].Value' --output text)
aws ec2 create-launch-template --launch-template-name $LAB-lt \
--launch-template-data "{\"ImageId\":\"$AMI\",\"InstanceType\":\"t3.micro\"}" >/dev/null
aws autoscaling create-auto-scaling-group --auto-scaling-group-name $LAB-asg \
--launch-template LaunchTemplateName=$LAB-lt,Version='$Latest' \
--min-size 2 --max-size 2 --desired-capacity 2 \
--vpc-zone-identifier "$SUB_A,$SUB_B"
sleep 60
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names $LAB-asg \
--query 'AutoScalingGroups[0].Instances[].{Id:InstanceId,AZ:AvailabilityZone,State:LifecycleState,Health:HealthStatus}' \
--output table
Q2. Record the two instances and their AZs. Did Auto Scaling place one in each? Quote the sentence from the Auto Scaling documentation that predicted this.
Now kill one:
VICTIM=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names $LAB-asg \
--query 'AutoScalingGroups[0].Instances[0].InstanceId' --output text)
echo "terminating $VICTIM"
aws ec2 terminate-instances --instance-ids $VICTIM >/dev/null
Watch (re-run every 30 seconds until it settles):
aws autoscaling describe-scaling-activities --auto-scaling-group-name $LAB-asg --max-items 5 \
--query 'Activities[].{Time:StartTime,Status:StatusCode,Cause:Cause}' --output table
Q3. Paste the Cause text for the replacement activity. In AWS's own words, what triggered it?
Q4. This group has no scaling policy at all. Write one sentence explaining what it is for, and name the requirement phrasing in an exam stem that this answers.
Q5. Now the max trap. Set desired to 4 and observe:
aws autoscaling set-desired-capacity --auto-scaling-group-name $LAB-asg --desired-capacity 4 2>&1 | tail -3
Paste the error. Which of min/max/desired stopped you, and what is the one-line fix?
This is the deliberate failure: you will cause uneven distribution and then diagnose it.
SG=$(aws ec2 create-security-group --group-name $LAB-sg --description "$LAB" \
--vpc-id $VPC --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id $SG \
--protocol tcp --port 80 --cidr 0.0.0.0/0 >/dev/null
ALB=$(aws elbv2 create-load-balancer --name $LAB-alb --type application \
--subnets $SUB_A $SUB_B --security-groups $SG \
--query 'LoadBalancers[0].LoadBalancerArn' --output text)
TG=$(aws elbv2 create-target-group --name $LAB-tg --protocol HTTP --port 80 \
--vpc-id $VPC --target-type instance --query 'TargetGroups[0].TargetGroupArn' --output text)
aws elbv2 create-listener --load-balancer-arn $ALB --protocol HTTP --port 80 \
--default-actions Type=forward,TargetGroupArn=$TG >/dev/null
Q6. Check the cross-zone setting on the target group, and on the load balancer:
aws elbv2 describe-target-group-attributes --target-group-arn $TG \
--query "Attributes[?contains(Key,'cross_zone')]"
aws elbv2 describe-load-balancer-attributes --load-balancer-arn $ALB \
--query "Attributes[?contains(Key,'cross_zone')]"
Record both. Lesson 3 flagged the per-type default as unverified — you have now verified it for your own account and load balancer type. Write down what you found, and what it means for the uneven-distribution scenario.
Q7. Attach the ASG to the target group and deliberately unbalance the AZs:
aws autoscaling attach-load-balancer-target-groups \
--auto-scaling-group-name $LAB-asg --target-group-arns $TG
# force everything into ONE AZ
aws autoscaling update-auto-scaling-group --auto-scaling-group-name $LAB-asg \
--vpc-zone-identifier "$SUB_A" --min-size 2 --max-size 2 --desired-capacity 2
sleep 90
aws elbv2 describe-target-health --target-group-arn $TG \
--query 'TargetHealthDescriptions[].{Id:Target.Id,AZ:Target.AvailabilityZone,State:TargetHealth.State,Reason:TargetHealth.Reason}' \
--output table
Record the table. All targets are now in one AZ, but the load balancer still has a subnet in two.
Q8. Resolve the ALB's DNS name and count the addresses:
DNS=$(aws elbv2 describe-load-balancers --load-balancer-arns $ALB \
--query 'LoadBalancers[0].DNSName' --output text)
dig +short $DNS
How many IP addresses came back, and how many AZs does the load balancer have enabled? Now read the NLB introduction quote from lesson 3 about DNS removal. Does what you observed match it? If it doesn't, say so — that is a legitimate finding, and note which load balancer type the quote was about.
Q9. Targets are unhealthy (nothing is listening on port 80 — we never installed a web server).
Explain why that is useful for this lab, and what the Reason field told you.
Q10. Given 2 targets in AZ-A and 8 in AZ-B with cross-zone disabled, compute the share of traffic each individual target receives. Show the arithmetic.
Q11. Name the two legitimate fixes for uneven distribution, and quote the AWS sentence that recommends the second one.
QURL=$(aws sqs create-queue --queue-name $LAB-q --query QueueUrl --output text)
DLQ=$(aws sqs create-queue --queue-name $LAB-dlq --query QueueUrl --output text)
DLQ_ARN=$(aws sqs get-queue-attributes --queue-url $DLQ \
--attribute-names QueueArn --query 'Attributes.QueueArn' --output text)
Q12. Wire up the dead-letter queue with a redrive policy of 2 receives, then confirm it:
aws sqs set-queue-attributes --queue-url $QURL --attributes \
"{\"RedrivePolicy\":\"{\\\"deadLetterTargetArn\\\":\\\"$DLQ_ARN\\\",\\\"maxReceiveCount\\\":\\\"2\\\"}\",\"VisibilityTimeout\":\"5\"}"
aws sqs get-queue-attributes --queue-url $QURL \
--attribute-names RedrivePolicy VisibilityTimeout --query Attributes
Record the output. What does maxReceiveCount of 2 mean in plain words?
Q13. Now be the poison message. Send one, receive it twice without deleting it, and find it in the DLQ:
aws sqs send-message --queue-url $QURL --message-body "poison" >/dev/null
for i in 1 2 3; do
echo "--- receive $i ---"
aws sqs receive-message --queue-url $QURL --wait-time-seconds 10 \
--query 'Messages[0].MessageId' --output text
sleep 6 # let the 5s visibility timeout expire
done
aws sqs receive-message --queue-url $DLQ --wait-time-seconds 10 \
--query 'Messages[0].Body' --output text
On which receive did the message stop coming back, and where did it end up? Explain what would have happened without a dead-letter queue, including how long it would have gone on.
Q14. You just reproduced duplicate delivery on purpose. State the design rule that follows, and the two configuration values that were in tension.
Q15. Fan-out. Create a topic, subscribe two queues, publish once:
TOPIC=$(aws sns create-topic --name $LAB-topic --query TopicArn --output text)
for N in a b; do
Q=$(aws sqs create-queue --queue-name $LAB-fan-$N --query QueueUrl --output text)
QA=$(aws sqs get-queue-attributes --queue-url $Q --attribute-names QueueArn \
--query 'Attributes.QueueArn' --output text)
aws sqs set-queue-attributes --queue-url $Q --attributes \
"{\"Policy\":\"{\\\"Version\\\":\\\"2012-10-17\\\",\\\"Statement\\\":[{\\\"Effect\\\":\\\"Allow\\\",\\\"Principal\\\":{\\\"Service\\\":\\\"sns.amazonaws.com\\\"},\\\"Action\\\":\\\"sqs:SendMessage\\\",\\\"Resource\\\":\\\"$QA\\\",\\\"Condition\\\":{\\\"ArnEquals\\\":{\\\"aws:SourceArn\\\":\\\"$TOPIC\\\"}}}]}\"}"
aws sns subscribe --topic-arn $TOPIC --protocol sqs --notification-endpoint $QA >/dev/null
eval "Q_$N=$Q"
done
aws sns publish --topic-arn $TOPIC --message "one event" >/dev/null
sleep 5
for N in a b; do
eval "Q=\$Q_$N"
echo -n "queue $N: "
aws sqs receive-message --queue-url $Q --wait-time-seconds 5 \
--query 'length(Messages)' --output text
done
Q16. How many copies of one published message exist? Now answer the design question: if queue b's
consumer were down for an hour, would the message survive? Would it have survived with a direct
SNS→Lambda subscription instead? Explain the difference in one sentence.
Q17. You had to attach a queue policy allowing sns.amazonaws.com to sqs:SendMessage. Which
lesson from Domain 1 does that demonstrate, and what would have happened without it?
No commands. These are the questions an interviewer or an exam stem will actually ask.
Q18. For each requirement, name the DR strategy and justify it in one sentence with a quote:
| # | Requirement |
|---|---|
| a | RPO 24 h, RTO 24 h, minimum cost |
| b | RPO 5 min, RTO 30 min, cost-conscious |
| c | RPO seconds, RTO minutes, cannot wait for deployment |
| d | RTO and RPO near zero, cost secondary |
Q19. Your DR plan uses Auto Scaling to reach production capacity in the recovery Region, and shifts traffic by changing Route 53 weights. Identify both control-plane dependencies, quote AWS on each, and give the data-plane alternative.
Q20. An engineer runs a destructive DELETE against production at 09:00. You replicate continuously
to a second Region with sub-second lag. State the time the problem reaches your DR Region, and name the
three controls that would have helped (one of them is from SAA1).
aws elbv2 delete-listener --listener-arn $(aws elbv2 describe-listeners \
--load-balancer-arn $ALB --query 'Listeners[0].ListenerArn' --output text) 2>/dev/null
aws elbv2 delete-load-balancer --load-balancer-arn $ALB
sleep 30
aws elbv2 delete-target-group --target-group-arn $TG
aws autoscaling update-auto-scaling-group --auto-scaling-group-name $LAB-asg \
--min-size 0 --max-size 0 --desired-capacity 0
aws autoscaling delete-auto-scaling-group --auto-scaling-group-name $LAB-asg --force-delete
aws ec2 delete-launch-template --launch-template-name $LAB-lt >/dev/null
for U in $QURL $DLQ $Q_a $Q_b; do aws sqs delete-queue --queue-url $U 2>/dev/null; done
aws sns delete-topic --topic-arn $TOPIC
sleep 60
aws ec2 delete-security-group --group-id $SG
Verify it — don't assume it:
aws elbv2 describe-load-balancers --query "LoadBalancers[?starts_with(LoadBalancerName,'$LAB')].LoadBalancerName"
aws autoscaling describe-auto-scaling-groups --query "AutoScalingGroups[?starts_with(AutoScalingGroupName,'$LAB')].AutoScalingGroupName"
aws sqs list-queues --queue-name-prefix $LAB
aws ec2 describe-instances --filters Name=tag:aws:autoscaling:groupName,Values=$LAB-asg \
Name=instance-state-name,Values=running --query 'Reservations[].Instances[].InstanceId'
All four should be empty. If the security group deletion failed, wait a few minutes — ENIs take time to detach — and retry.
sleep 6 must exceed the 5-second visibility timeout. If
someone's shell lags, have them re-run rather than debug it.