AWS Training
Modules Listen All tracks

← Design High-Performing Architectures

SAA3 Lab — measure the bottleneck

You will need: AWS CLI v2, jq, python3, and permission to create EBS volumes, IAM roles, Lambda functions, Kinesis streams, S3 buckets and Athena queries.

Write your answers down. Questions are numbered Q1…Q20.

⚠️ Safety: everything here is cheap. The one thing that bills by the hour is the Kinesis stream. Run the teardown today, and run the verification block after it.

export LAB=saa3-$(date +%Y%m%d)-$RANDOM
export REGION=$(aws configure get region)
export ACCT=$(aws sts get-caller-identity --query Account --output text)
export AZ=$(aws ec2 describe-availability-zones --query 'AvailabilityZones[0].ZoneName' --output text)
echo "LAB=$LAB REGION=$REGION AZ=$AZ"

Part 1 — EBS: the limits are real (20 min)

Create two small volumes that you will never attach:

GP2=$(aws ec2 create-volume --availability-zone $AZ --volume-type gp2 --size 10 \
  --tag-specifications "ResourceType=volume,Tags=[{Key=lab,Value=$LAB}]" --query VolumeId --output text)
GP3=$(aws ec2 create-volume --availability-zone $AZ --volume-type gp3 --size 10 \
  --tag-specifications "ResourceType=volume,Tags=[{Key=lab,Value=$LAB}]" --query VolumeId --output text)
sleep 10
aws ec2 describe-volumes --volume-ids $GP2 $GP3 \
  --query 'Volumes[].{Id:VolumeId,Type:VolumeType,Size:Size,Iops:Iops,Throughput:Throughput}' --output table

Q1. Record the IOPS for each. Explain the gp2 number using the "3 IOPS per GiB" rule and the minimum from lesson 1. Explain the gp3 number using the baseline sentence.

Q2. Now test gp3's "500 IOPS per GiB" ratio on a 10 GiB volume:

aws ec2 modify-volume --volume-id $GP3 --iops 4000 --query 'VolumeModification.TargetIops'
aws ec2 modify-volume --volume-id $GP3 --iops 6000 2>&1 | tail -2

Which call succeeded and which failed? Compute the maximum IOPS this volume can have, and state what you'd change to reach 80,000.

Q3. Now try to create an HDD volume below its documented minimum:

aws ec2 create-volume --availability-zone $AZ --volume-type st1 --size 10 2>&1 | tail -2

Paste the error. Which lesson 1 number predicted it? (If it unexpectedly succeeds, record that as a finding and delete the volume it created.)

Q4. Without running anything: a stem says "put the operating system on an sc1 volume to save money". Quote the line from the EBS volume-types page that eliminates it.


Part 2 — Instance types: read the name, then check it (15 min)

aws ec2 describe-instance-types \
  --instance-types t3.micro m7i.large c7gd.large r7g.large \
  --query 'InstanceTypes[].{Type:InstanceType,vCPU:VCpuInfo.DefaultVCpus,MemMiB:MemoryInfo.SizeInMiB,LocalDisk:InstanceStorageSupported,Burstable:BurstablePerformanceSupported,Net:NetworkInfo.NetworkPerformance,ENA:NetworkInfo.EnaSupport,EFA:NetworkInfo.EfaSupported}' \
  --output table

Q5. Before looking at the output, decode each name (series, generation, options). Then check: does LocalDisk match the d option letter? Does Burstable match the T series?

Q6. Compare memory per vCPU for c7gd.large and r7g.large. Which workload from lesson 2's choosing table would you put on each?

Q7. Is ENA supported on all four? Quote the enhanced-networking sentence about Nitro instances that predicts this.


Part 3 — Lambda: memory is the CPU knob (25 min)

cat > /tmp/$LAB-trust.json <<'EOF'
{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":{"Service":"lambda.amazonaws.com"},"Action":"sts:AssumeRole"}]}
EOF
ROLE_ARN=$(aws iam create-role --role-name $LAB-role \
  --assume-role-policy-document file:///tmp/$LAB-trust.json --query Role.Arn --output text)
aws iam attach-role-policy --role-name $LAB-role \
  --policy-arn arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole

cat > /tmp/lambda_function.py <<'EOF'
import hashlib
def handler(event, context):
    h = b"x"
    for _ in range(300000):
        h = hashlib.sha256(h).digest()
    return {"done": True}
EOF
(cd /tmp && zip -q $LAB.zip lambda_function.py)
sleep 10   # IAM role propagation
aws lambda create-function --function-name $LAB-fn --runtime python3.12 \
  --role $ROLE_ARN --handler lambda_function.handler \
  --zip-file fileb:///tmp/$LAB.zip --memory-size 128 --timeout 60 >/dev/null
aws lambda wait function-active-v2 --function-name $LAB-fn

Invoke it at three memory sizes:

for MEM in 128 1769 3538; do
  aws lambda update-function-configuration --function-name $LAB-fn --memory-size $MEM >/dev/null
  aws lambda wait function-updated-v2 --function-name $LAB-fn
  for i in 1 2; do
    aws lambda invoke --function-name $LAB-fn --log-type Tail /tmp/$LAB-out.json \
      --query LogResult --output text | base64 --decode | grep REPORT
  done
done

Q8. Record Duration, Billed Duration and Memory Size for each (use the second invocation of each pair — the first may include a cold start). How much faster was 1,769 MB than 128 MB?

Q9. Lambda bills per request plus GB-seconds (SAA2 lesson 4), so compute cost scales with memory × duration. Compute memory (GB) × billed duration (s) for 128 MB and 1,769 MB. Which was cheaper for this CPU-bound function? Explain using the "CPU power in proportion to the amount of memory" sentence.

Q10. Did 3,538 MB (two vCPU-equivalents) halve the duration again? If not, why not? (Hint: is this function's loop parallel?) What does that tell you about "just add memory" as a rule?


Part 4 — Kinesis: shards and partition keys (20 min)

aws kinesis create-stream --stream-name $LAB-stream --shard-count 2
aws kinesis wait stream-exists --stream-name $LAB-stream
aws kinesis describe-stream-summary --stream-name $LAB-stream \
  --query 'StreamDescriptionSummary.{Mode:StreamModeDetails.StreamMode,Shards:OpenShardCount,RetentionH:RetentionPeriodHours}'

Q11. Record mode, shard count and retention. Does the retention match the documented default?

Q12. Send ten records with the same partition key, then ten with different keys:

for i in $(seq 1 10); do
  aws kinesis put-record --stream-name $LAB-stream --partition-key user-42 \
    --data "click $i" --cli-binary-format raw-in-base64-out --query ShardId --output text
done | sort | uniq -c
for i in $(seq 1 10); do
  aws kinesis put-record --stream-name $LAB-stream --partition-key user-$RANDOM \
    --data "click $i" --cli-binary-format raw-in-base64-out --query ShardId --output text
done | sort | uniq -c

Paste both counts. Explain the first using the MD5 partition-key sentence. What production symptom would the first pattern cause at scale, and what is the fix?

Q13. Increase retention to 7 days and confirm:

aws kinesis increase-stream-retention-period --stream-name $LAB-stream --retention-period-hours 168
sleep 5
aws kinesis describe-stream-summary --stream-name $LAB-stream \
  --query StreamDescriptionSummary.RetentionPeriodHours

What is the maximum you could set, and what does the key-concepts page say about the charge for going above 24 hours?

Q14. Using the provisioned-mode formula: 2,500 records/s at 3 KB each, three consumers on shared throughput. How many shards? Show the arithmetic. How would enhanced fan-out change the answer?


Part 5 — Athena: CSV vs Parquet, measured (25 min)

aws s3 mb s3://$LAB-bucket >/dev/null
python3 - <<'EOF' > /tmp/sales.csv
import random
random.seed(1)
print("order_id,region,product,qty,price")
for i in range(200000):
    print(f"{i},{random.choice(['us','eu','ap'])},p{random.randint(1,500)},{random.randint(1,9)},{random.randint(100,9999)/100}")
EOF
aws s3 cp /tmp/sales.csv s3://$LAB-bucket/csv/sales.csv >/dev/null

athena() {
  local QID=$(aws athena start-query-execution --query-string "$1" \
    --result-configuration OutputLocation=s3://$LAB-bucket/results/ --query QueryExecutionId --output text)
  while true; do
    local S=$(aws athena get-query-execution --query-execution-id $QID --query QueryExecution.Status.State --output text)
    [ "$S" = "SUCCEEDED" ] || [ "$S" = "FAILED" ] || [ "$S" = "CANCELLED" ] && break; sleep 2
  done
  aws athena get-query-execution --query-execution-id $QID \
    --query 'QueryExecution.{State:Status.State,Scanned:Statistics.DataScannedInBytes,Ms:Statistics.EngineExecutionTimeInMillis,Why:Status.StateChangeReason}'
}
DB=$(echo $LAB | tr '-' '_')
athena "CREATE DATABASE $DB"
athena "CREATE EXTERNAL TABLE $DB.sales_csv (order_id bigint, region string, product string, qty int, price double) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' LOCATION 's3://$LAB-bucket/csv/' TBLPROPERTIES ('skip.header.line.count'='1')"
athena "CREATE TABLE $DB.sales_parquet WITH (format='PARQUET', external_location='s3://$LAB-bucket/parquet/') AS SELECT * FROM $DB.sales_csv"

Q15. You just did the guide's "Transforming data between formats (for example, .csv to .parquet)" with one statement. Which of the three conversion methods in lesson 6 was it?

Q16. Run the same single-column aggregate against both tables:

athena "SELECT sum(price) FROM $DB.sales_csv"
athena "SELECT sum(price) FROM $DB.sales_parquet"

Record Scanned for each. Compute the ratio. Quote the lesson 6 bullet that explains it.

Q17. Compare aws s3 ls s3://$LAB-bucket/csv/ --summarize with aws s3 ls s3://$LAB-bucket/parquet/ --summarize. Which of Parquet's three properties does the size difference demonstrate?


Part 6 — Reading the plan, not the console (15 min, no writes)

Q18. For each requirement, name the answer and quote the sentence or number that decides it:

# Requirement
a Partners must allowlist two fixed IPs; the app runs in two Regions
b Dedicated 10 Gbps link to on-premises, needed in three days
c HPC nodes exchange MPI traffic; lowest latency
d Application subnet is a /28; the team is moving 12 tasks to Fargate

Q19. A Lambda-backed API is failing with RDS connection errors under load, and the team proposes adding two read replicas. Explain why that won't fix it, what will, and one limit of the fix.

Q20. A DynamoDB table needs microsecond reads for a product catalogue that changes once a day. Another table records trades and every read must reflect the latest write. Which one gets DAX, and why not the other? Quote AWS.


Teardown — today

aws kinesis delete-stream --stream-name $LAB-stream
aws lambda delete-function --function-name $LAB-fn
aws iam detach-role-policy --role-name $LAB-role \
  --policy-arn arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole
aws iam delete-role --role-name $LAB-role
aws ec2 delete-volume --volume-id $GP2
aws ec2 delete-volume --volume-id $GP3
athena "DROP TABLE $DB.sales_parquet"
athena "DROP TABLE $DB.sales_csv"
athena "DROP DATABASE $DB"
aws s3 rb s3://$LAB-bucket --force
rm -f /tmp/$LAB* /tmp/lambda_function.py /tmp/sales.csv

Verify it — don't assume it:

aws kinesis list-streams --query "StreamNames[?starts_with(@,'$LAB')]"
aws lambda list-functions --query "Functions[?starts_with(FunctionName,'$LAB')].FunctionName"
aws ec2 describe-volumes --filters Name=tag:lab,Values=$LAB --query 'Volumes[].VolumeId'
aws s3 ls | grep $LAB
aws iam get-role --role-name $LAB-role 2>&1 | tail -1

The first four should be empty; the last should say the role cannot be found. If Part 1's Q3 created a volume unexpectedly, delete it by ID as well — it won't carry the lab tag.


Done when you can

Facilitator notes