AWS Training
Modules Listen All tracks

← Design Secure Architectures

SAA1 Lab — Design Secure Architectures

You will need: an AWS account you can create IAM roles and KMS keys in, the AWS CLI v2 configured, and jq. Part 4 needs an AWS Organizations management account — if you don't have one, read it and skip the commands; everything else runs in a single account.

Write your answers down. The point of this lab is not to run commands — it's to produce evidence and then explain it. Questions are numbered Q1…Q18. A checklist and facilitator notes are at the end.

⚠️ Two safety rules, and they are not optional.

  1. Do not enable S3 Object Lock in COMPLIANCE mode in this lab. Per AWS: "The only way to delete an object under the compliance mode before its retention date expires is to delete the associated AWS account." Part 5 uses GOVERNANCE mode, which is designed to be reversible.
  2. Do not detach FullAWSAccess from anything. Part 4 attaches an additional deny SCP to a test OU only. Detaching FullAWSAccess would break every member account under that level.

Set up a scratch prefix so teardown is easy:

export LAB=saa1-$(date +%Y%m%d)-$RANDOM
export ACCT=$(aws sts get-caller-identity --query Account --output text)
export REGION=$(aws configure get region)
echo "LAB=$LAB ACCT=$ACCT REGION=$REGION"

Part 1 — Prove the union (25 min)

Task statement 1.1. You are going to demonstrate that a resource policy alone grants access in the same account.

# A bucket, and a role with NO S3 permissions at all
aws s3api create-bucket --bucket $LAB-union --region $REGION \
  $( [ "$REGION" = "us-east-1" ] || echo "--create-bucket-configuration LocationConstraint=$REGION" )

cat > /tmp/$LAB-trust.json <<JSON
{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
 "Principal":{"AWS":"arn:aws:iam::$ACCT:root"},"Action":"sts:AssumeRole"}]}
JSON

aws iam create-role --role-name $LAB-nothing \
  --assume-role-policy-document file:///tmp/$LAB-trust.json

# Deliberately attach NO permissions policy. Confirm that:
aws iam list-attached-role-policies --role-name $LAB-nothing
aws iam list-role-policies --role-name $LAB-nothing

Q1. Record both outputs. What permissions does this role have right now, and which of the three evaluation outcomes from lesson 1 applies?

Now assume it and try to read the bucket:

CREDS=$(aws sts assume-role --role-arn arn:aws:iam::$ACCT:role/$LAB-nothing \
  --role-session-name test1 --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo $CREDS | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo $CREDS | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo $CREDS | jq -r .SessionToken)

aws sts get-caller-identity          # confirm you are the role
aws s3api list-objects-v2 --bucket $LAB-union   # expect AccessDenied
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN

Q2. Paste the exact error. Which policy denied it — or did nothing allow it? Quote the difference.

Now add a bucket policy naming the role, and nothing else:

cat > /tmp/$LAB-bucket.json <<JSON
{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
 "Principal":{"AWS":"arn:aws:iam::$ACCT:role/$LAB-nothing"},
 "Action":["s3:ListBucket"],"Resource":"arn:aws:s3:::$LAB-union"}]}
JSON
aws s3api put-bucket-policy --bucket $LAB-union --policy file:///tmp/$LAB-bucket.json

Re-assume the role and list the bucket again.

Q3. It now works. The role still has no identity-based policy. Write one sentence explaining this using the word union, and one sentence explaining why the same thing would fail if the bucket were in a different account.

Q4. Add "Effect":"Deny" for s3:ListBucket to the role as an inline policy, keep the bucket policy, and retry. Record the result. Which rule from lesson 1 did you just demonstrate?


Part 2 — Break role chaining on purpose (20 min)

This is the deliberate failure. You will cause it, diagnose it, and then explain it.

# Role A can assume Role B; both have a 12-hour max session
aws iam update-role --role-name $LAB-nothing --max-session-duration 43200

cat > /tmp/$LAB-trustB.json <<JSON
{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
 "Principal":{"AWS":"arn:aws:iam::$ACCT:role/$LAB-nothing"},"Action":"sts:AssumeRole"}]}
JSON
aws iam create-role --role-name $LAB-chained \
  --assume-role-policy-document file:///tmp/$LAB-trustB.json --max-session-duration 43200

cat > /tmp/$LAB-canassume.json <<JSON
{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
 "Action":"sts:AssumeRole","Resource":"arn:aws:iam::$ACCT:role/$LAB-chained"}]}
JSON
aws iam put-role-policy --role-name $LAB-nothing \
  --policy-name can-chain --policy-document file:///tmp/$LAB-canassume.json

Direct assumption with a long duration — this should work:

aws sts assume-role --role-arn arn:aws:iam::$ACCT:role/$LAB-chained \
  --role-session-name direct --duration-seconds 43200 \
  --query 'Credentials.Expiration' --output text

Q5. Record the expiration. Roughly how far in the future is it?

Now chain, with the same duration:

CREDS=$(aws sts assume-role --role-arn arn:aws:iam::$ACCT:role/$LAB-nothing \
  --role-session-name hop1 --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo $CREDS | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo $CREDS | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo $CREDS | jq -r .SessionToken)

# THIS IS THE DELIBERATE FAILURE
aws sts assume-role --role-arn arn:aws:iam::$ACCT:role/$LAB-chained \
  --role-session-name hop2 --duration-seconds 43200

Q6. Paste the exact error message. Both roles are configured for 12 hours. Why did this fail?

Q7. Find the largest --duration-seconds that succeeds from the chained session. State the number and the documented rule it matches.

# then clean up your shell
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN

Q8. A CI pipeline assumes a deploy role, then assumes a prod role, and needs an 8-hour session for a long migration. Write the two-sentence design recommendation you would give that team.


Part 3 — Lock yourself out of a KMS key (30 min)

Task statement 1.3. This part demonstrates the key-policy inversion. Read the whole part before running anything — you will deliberately create a key you cannot use, then recover it via the one route that works.

KEYID=$(aws kms create-key --description "$LAB scratch" \
  --query KeyMetadata.KeyId --output text)
aws kms create-alias --alias-name alias/$LAB --target-key-id $KEYID
aws kms get-key-policy --key-id $KEYID --policy-name default \
  --output text --query Policy | jq .

Q9. In the default key policy, find the statement whose Action is kms:* with the account root as principal. Quote it. What does the AWS documentation say this statement is for?

Confirm you can use the key, then confirm rotation is off, then turn it on:

echo -n "hello" > /tmp/$LAB-pt
aws kms encrypt --key-id alias/$LAB --plaintext fileb:///tmp/$LAB-pt \
  --query CiphertextBlob --output text > /tmp/$LAB-ct

aws kms get-key-rotation-status --key-id $KEYID
aws kms enable-key-rotation --key-id $KEYID
aws kms get-key-rotation-status --key-id $KEYID

Q10. Record both rotation statuses. What is the RotationPeriodInDays, and what did you have to specify to get that value?

Now perform an on-demand rotation and check what changed:

aws kms rotate-key-on-demand --key-id $KEYID
aws kms list-key-rotations --key-id $KEYID
aws kms describe-key --key-id $KEYID --query 'KeyMetadata.{Id:KeyId,Arn:Arn}'

# decrypt the ciphertext you made BEFORE the rotation
aws kms decrypt --ciphertext-blob fileb://<(base64 -d /tmp/$LAB-ct) \
  --key-id alias/$LAB --query Plaintext --output text | base64 -d; echo

Q11. The key ID and ARN — did they change? Did the old ciphertext still decrypt? Quote the documentation sentence that explains both answers at once.

Q12. Your CISO says "we rotated the key, so the backup that leaked last year is now safe." Write the two-sentence correction, citing the documentation.

Now the lockout. Replace the key policy with one that omits the IAM-enabling statement:

cat > /tmp/$LAB-badkey.json <<JSON
{"Version":"2012-10-17","Id":"lockout-demo","Statement":[
 {"Sid":"OnlyOneSpecificRole","Effect":"Allow",
  "Principal":{"AWS":"arn:aws:iam::$ACCT:role/$LAB-chained"},
  "Action":["kms:Decrypt"],"Resource":"*"}]}
JSON
aws kms put-key-policy --key-id $KEYID --policy-name default \
  --policy file:///tmp/$LAB-badkey.json

# You are an admin. Try to use your own key.
aws kms encrypt --key-id $KEYID --plaintext fileb:///tmp/$LAB-pt

Q13. Paste the error. You are an account administrator with kms:* in your IAM policy. Explain, in one sentence quoting the documentation, why your IAM policy had no effect.

Q14. Who can still change this key policy, and why? (Try it — put-key-policy back to the default, using the policy you saved in Q9. Record whether it worked and reason about why.)

⚠️ If you cannot restore the policy, you have genuinely locked the key. Schedule it for deletion (aws kms schedule-key-deletion --key-id $KEYID --pending-window-in-days 7) and note in your answers that this was the only remaining option. That is itself the lesson.


Part 4 — An SCP guardrail (25 min) — management account required

# Confirm you are in an organization with all features
aws organizations describe-organization --query 'Organization.FeatureSet'
ROOT=$(aws organizations list-roots --query 'Roots[0].Id' --output text)
OU=$(aws organizations create-organizational-unit --parent-id $ROOT \
  --name $LAB-sandbox --query 'OrganizationalUnit.Id' --output text)
aws organizations list-policies-for-target --target-id $OU --filter SERVICE_CONTROL_POLICY

Q15. What policy is already attached to the brand-new OU, and who attached it? What would happen to accounts in this OU if you detached it and attached nothing? Quote the documentation. Do not detach it.

Create and attach an additional deny SCP:

cat > /tmp/$LAB-scp.json <<'JSON'
{"Version":"2012-10-17","Statement":[
 {"Sid":"DenyLeaveOrg","Effect":"Deny",
  "Action":"organizations:LeaveOrganization","Resource":"*"}]}
JSON
POL=$(aws organizations create-policy --name $LAB-deny-leave \
  --type SERVICE_CONTROL_POLICY --description "lab" \
  --content file:///tmp/$LAB-scp.json --query 'Policy.PolicySummary.Id' --output text)
aws organizations attach-policy --policy-id $POL --target-id $OU
aws organizations list-policies-for-target --target-id $OU --filter SERVICE_CONTROL_POLICY

Q16. Two SCPs are now attached. Using lesson 2's allow/deny asymmetry, state in one sentence why both are required for accounts in this OU to work normally and be prevented from leaving.

Q17. Now test the size rule. Take the policy above, add 10,300 characters of white space between JSON elements, and try to create it via the CLI. Then paste the same content into the console's policy editor. Record what happens in each case and quote the documentation sentence that explains the difference.


Part 5 — Permissions are not immutability (30 min)

This is the part that matters most. Task statement 1.3.

aws s3api create-bucket --bucket $LAB-lock --region $REGION \
  $( [ "$REGION" = "us-east-1" ] || echo "--create-bucket-configuration LocationConstraint=$REGION" ) \
  --object-lock-enabled-for-bucket

aws s3api get-bucket-versioning --bucket $LAB-lock
echo "keep me" > /tmp/$LAB-obj
aws s3api put-object --bucket $LAB-lock --key important.txt --body /tmp/$LAB-obj

Q18a. You passed --object-lock-enabled-for-bucket and did not ask for versioning. What does get-bucket-versioning return, and which documented requirement explains it?

Apply GOVERNANCE mode retention (not compliance — see the safety rules):

VID=$(aws s3api list-object-versions --bucket $LAB-lock --prefix important.txt \
  --query 'Versions[0].VersionId' --output text)
UNTIL=$(date -u -v+2d +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '+2 days' +%Y-%m-%dT%H:%M:%SZ)

aws s3api put-object-retention --bucket $LAB-lock --key important.txt \
  --version-id $VID --retention "Mode=GOVERNANCE,RetainUntilDate=$UNTIL"
aws s3api get-object-retention --bucket $LAB-lock --key important.txt --version-id $VID

Now the two deletes. Predict each result before you run it.

# (a) permanent delete — specifies the version ID
aws s3api delete-object --bucket $LAB-lock --key important.txt --version-id $VID

# (b) simple delete — no version ID
aws s3api delete-object --bucket $LAB-lock --key important.txt
aws s3api list-object-versions --bucket $LAB-lock --prefix important.txt

Q18b. Record both results and both HTTP behaviours. Why did (b) "succeed"? Where is the data?

Q18c. Now the bypass. Run the permanent delete again with --bypass-governance-retention. Did it work? Which two things does the documentation say you need, and which one did the CLI flag supply? What does the S3 console do differently, and why does that weaken governance mode as a control?

Finally, the headline experiment. Deny everything, then add a lifecycle rule:

cat > /tmp/$LAB-denyall.json <<JSON
{"Version":"2012-10-17","Statement":[{"Sid":"DenyAll","Effect":"Deny",
 "Principal":"*","Action":"s3:*","Resource":[
 "arn:aws:s3:::$LAB-lock","arn:aws:s3:::$LAB-lock/*"]}]}
JSON
aws s3api put-bucket-policy --bucket $LAB-lock --policy file:///tmp/$LAB-denyall.json

# confirm the deny bites for a principal
aws s3api list-objects-v2 --bucket $LAB-lock        # expect AccessDenied

Q18d. Read the Important box on Managing the lifecycle of objects and quote it. Given a bucket policy that denies all actions to all principals, would a lifecycle expiration rule still delete objects in this bucket? Write down your answer and the quote that proves it.

Q18e. A stakeholder says: "we've denied s3:DeleteObject for everyone, so the data can't be deleted." Write the three-sentence correction, name the control that actually provides the guarantee, and say what the trade-off is.

Note on Q18d: you are not asked to run a lifecycle expiration against your own locked object — the documented behaviour is what's being tested, and Object Lock would in any case protect the locked version. If you want to observe lifecycle expiry, do it in a separate un-locked bucket with a 1-day expiration rule and come back tomorrow.


Teardown — do this today

# Part 5: remove the deny policy first, or you cannot clean up
aws s3api delete-bucket-policy --bucket $LAB-lock
# Governance-mode objects need the bypass to remove
for V in $(aws s3api list-object-versions --bucket $LAB-lock \
    --query 'Versions[].VersionId' --output text); do
  aws s3api delete-object --bucket $LAB-lock --key important.txt \
    --version-id $V --bypass-governance-retention
done
for V in $(aws s3api list-object-versions --bucket $LAB-lock \
    --query 'DeleteMarkers[].VersionId' --output text); do
  aws s3api delete-object --bucket $LAB-lock --key important.txt --version-id $V
done
aws s3api delete-bucket --bucket $LAB-lock

# Part 1
aws s3api delete-bucket-policy --bucket $LAB-union
aws s3api delete-bucket --bucket $LAB-union

# Parts 1–2
aws iam delete-role-policy --role-name $LAB-nothing --policy-name can-chain
aws iam delete-role --role-name $LAB-nothing
aws iam delete-role --role-name $LAB-chained

# Part 3 — 7 days is the minimum pending window; the key bills until it is gone
aws kms delete-alias --alias-name alias/$LAB
aws kms schedule-key-deletion --key-id $KEYID --pending-window-in-days 7

# Part 4
aws organizations detach-policy --policy-id $POL --target-id $OU
aws organizations delete-policy --policy-id $POL
aws organizations delete-organizational-unit --organizational-unit-id $OU

rm -f /tmp/$LAB-*

Verify the teardown rather than assuming it:

aws s3api list-buckets --query "Buckets[?starts_with(Name,'$LAB')].Name"
aws iam list-roles --query "Roles[?starts_with(RoleName,'$LAB')].RoleName"
aws kms describe-key --key-id $KEYID --query 'KeyMetadata.KeyState'

All three should be empty / PendingDeletion. If they aren't, you are still paying.


Done when you can

Facilitator notes