AWS Training
Modules Listen All tracks

← Design Secure Architectures

SAA1 Interview — AWS security architecture, spoken

The quiz tests recall. This file rehearses speech. Read the answers out loud. If one takes more than 90 seconds to say, it's too long — split it.

These are questions the subject matter justifies for a solutions architect or cloud engineer role. No claims here about what any particular company asks.


Warm-up

"Walk me through how AWS decides whether a request is allowed."

Three outcomes. If any policy that applies has an explicit deny, it's denied — nothing overrides that. If something allows it and nothing denies it and every ceiling permits it, it's allowed. Otherwise it's a default deny, because silence is a no.

The part people get wrong is how the policy types combine. Identity and resource policies in the same account form a union — either one is enough. Permissions boundaries, service control policies and session policies all intersect — they're ceilings, so they all have to allow. I remember it as one union, three intersections.

"What's the difference between a security group and a network ACL?"

Three things, really. A security group is at the instance level and a network ACL is at the subnet level. Security groups only have allow rules; network ACLs have allow and deny. And security groups are stateful, so return traffic is automatic, where network ACLs are stateless and you have to allow the response separately.

In practice I use security groups as the primary control and network ACLs for two specific jobs: blocking something, since only they can deny, and as a subnet-level backstop for an instance that got launched with the wrong security group.

"When would you use a NAT gateway?"

When instances in a private subnet need outbound access and nothing should be able to reach them inbound. Connections have to be initiated from inside the VPC — it's never an inbound path. So if somebody's asking how to expose a private instance to the internet, a NAT gateway isn't the answer, a load balancer is.

For IPv6 the equivalent is an egress-only internet gateway.

"Cognito user pool or identity pool?"

A user pool is a user directory and an identity provider. It authenticates people and hands back tokens.

An identity pool is a credentials broker. It takes proof of authentication and hands back temporary AWS credentials from STS.

So the test I use is: does the client need to call an AWS API directly? If the browser or the mobile app is talking to S3 or DynamoDB itself, I need an identity pool, because that's the only thing that produces AWS credentials. If it only talks to my own API, a user pool is enough.

"What does a service control policy do?"

It sets the maximum available permissions for principals in member accounts. It never grants anything — you still need real IAM policies underneath.

And two exceptions matter architecturally: it doesn't apply to the management account, and it doesn't apply to service-linked roles. Which is why I'd never run workloads in the management account — they'd sit above the entire guardrail model.

"Is S3 data encrypted at rest by default?"

Yes. Server-side encryption with S3-managed keys is applied to every bucket, and every new upload has been encrypted since January 2023. So "turn on encryption" isn't really a task any more.

The real question is which key and who controls it — S3-managed, KMS, dual-layer KMS, or customer-provided. That's a permissions and audit decision rather than a cryptography one.


Depth

"Explain envelope encryption, and why KMS doesn't just encrypt my file."

The key material inside KMS lives on a hardware security module and is designed never to leave it in plaintext. So KMS can't be in the data path for large objects — it'd be a bottleneck and the key would have to travel.

Instead you ask KMS for a data key. It gives you the same key twice: once in plaintext, once encrypted under your KMS key. You encrypt your data locally with the plaintext copy, discard that copy, and store the encrypted copy next to the ciphertext. To read the data you send the encrypted key back to KMS and it decrypts it on an HSM.

That's how you encrypt a terabyte with a service whose direct encrypt call only handles small payloads.

The follow-up they ask next: "So if someone steals the encrypted data key, what protects you?" The fact that it can only be decrypted on an HSM, via an authorised call to KMS. So the protection is the key policy, not the cryptography — which is the honest answer and the one that shows you understand where the boundary actually is.

"Why might an account administrator be unable to use a KMS key in their own account?"

Because KMS inverts the usual rule. Every KMS key must have exactly one key policy, and unless the key policy explicitly allows it, an IAM policy can't grant access to that key. The default key policy includes a statement that enables IAM policies, which is why nobody notices — but if someone writes a tight key policy and leaves that statement out, every IAM admin in the account loses access.

And it's not recoverable by an admin, because root has no inherent permission to a key either. That's how accounts permanently lock themselves out of their own keys.

The follow-up they ask next: "Can you deny access to a key with an IAM policy alone?" Yes — deny works without the key policy's cooperation. It's only allow that needs it. That asymmetry is worth knowing because it's how you add a guardrail to a key you don't own the policy for.

"We rotate our KMS keys annually. A backup from last year has leaked. Are we covered?"

No, and this is worth being direct about. Rotation changes only the current key material, for new encryption operations. AWS retains every previous version and automatically picks the right one on decrypt. So that old ciphertext decrypts exactly as it always did.

The documentation says it explicitly — key rotation will not mitigate the effect of a compromised data key, and it doesn't re-encrypt anything.

If the requirement is that old ciphertext becomes unreadable, rotation is the wrong control. You either re-encrypt the data — S3 Batch Operations for a bucket — or you destroy the key and accept that you lose access too.

The follow-up they ask next: "So what is rotation actually for?" Mostly compliance obligations, and limiting the amount of data protected by any one piece of key material going forward. AWS is candid that KMS keys are usually wrapping keys and are almost never used enough to risk key exhaustion.

"Our bucket policy denies s3:DeleteObject for every principal. Is the data safe from deletion?"

No. And this is the one I'd flag in a design review.

A bucket policy governs principals making requests. A lifecycle rule isn't a principal making a request — it's the service acting on its own configuration. The S3 documentation says that even if your bucket policy denies all actions for all principals, the lifecycle configuration still functions as normal.

So a misconfigured expiration rule deletes your data and the deny does nothing. If the requirement is genuinely "cannot be deleted", the control is S3 Object Lock, not a permission.

The follow-up they ask next: "Then which Object Lock mode?" Depends whether you want an escape hatch. Governance mode lets someone with the bypass permission delete. Compliance mode means nobody can, including root — and the documented way to delete before the retention date is to close the AWS account. So I'd want compliance mode signed off by someone who can own a seven-year commitment.

"Walk me through securing a three-tier web app in a VPC."

Public subnets for the load balancer, private subnets for the application, private subnets with no outbound route for the database. Then chain the security groups by reference rather than by CIDR: the load balancer's group allows 443 from the internet, the app tier's group allows traffic only from the load balancer's group, and the database's group allows its port only from the app tier's group.

The reason I use group references rather than CIDR ranges is that they follow instances when the app autoscales. A CIDR range doesn't.

Then network ACLs as a coarse subnet backstop, VPC endpoints so S3 and DynamoDB traffic never leaves the AWS network, and flow logs so I can actually see what happened.

The follow-up they ask next: "You've put an inspection appliance between two tiers. Anything break?" Yes — security group referencing. If traffic is routed through a middlebox, referencing the other instance's security group doesn't allow the traffic. You have to use the private IP or the subnet CIDR instead. It's a documented limitation and it surprises people mid-migration.

"How would you give a third-party vendor access to our account?"

A cross-account IAM role, never an IAM user with access keys. Their account is the principal in the trust policy, and I'd require an external ID.

The external ID matters because trusting an account trusts everyone in it. The vendor's own documentation usually gives you one; if they don't, I'd ask, because without it any principal in their account can assume my role — that's the confused deputy problem.

I'd also scope the permissions policy to exactly what they need, add a condition requiring MFA if they're doing anything privileged, and monitor sts:SourceIdentity in CloudTrail so I can tell who on their side actually did something.

The follow-up they ask next: "What if they say they can't support an external ID?" Then I'd compensate — narrow the permissions hard, add IP conditions if they publish egress ranges, and make the gap explicit in the risk register. I wouldn't just drop the requirement quietly.

"Which of these is a cost decision and which is an engineering decision: the Organizations account limit, the SCP size limit, and the KMS key charge?"

The account limit is a request — it defaults to ten, it's adjustable, and only the management account can ask. So it's a paperwork decision, and the answer is to file the request early rather than design around it.

The SCP size limit is 10,240 characters and it's a hard limit. That's an engineering decision — you consolidate policies or restructure your OUs. There's a nasty edge there too: the console strips white space before saving and the CLI doesn't, so the same policy can pass by hand and fail in CI.

The KMS key charge is genuinely a cost decision. A customer managed key has a monthly fee plus per-use charges; an AWS managed key has no monthly fee; an AWS owned key is free. But you can only share cross-account with a customer managed key, and you can't audit an AWS owned key at all. So it's cost versus control, and the requirement decides it, not the price.


Design

"Design data protection for a financial services firm. Seven-year retention, regulator can demand proof, and we've just been through a ransomware scare. Assume AWS."

Let me get the requirements straight first, because they pull in different directions. Seven-year immutable retention. Demonstrable evidence for a regulator. And resilience against an attacker who gets administrative credentials in production. That last one changes the design more than the other two.

For the S3 data: versioning, then Object Lock. I'd want compliance mode for the regulated records, because that's the only mode where root can't delete — and the documentation is blunt that the only way out early is closing the account, so that's a decision someone senior signs. For anything where I'm not certain of the retain-until date, I'd use a legal hold or variable retention instead, so I'm not committing to a date I'll regret.

For everything that isn't S3 — the databases, the EBS volumes, EFS — AWS Backup with backup plans assigned by tag, so new resources are covered without anyone remembering to add them. Vault Lock on the vaults for the same WORM guarantee.

Now the ransomware requirement, which is where it gets interesting. Backups in the production account are only as safe as production. So: cross-account copies fanned in to a dedicated repository account that production has no access to, inside Organizations. And I'd use full AWS Backup management so the backups are encrypted with the vault's KMS key rather than the source resource's key — otherwise an attacker who can disable the source key can make the backups useless without ever touching them.

The trade-off I'd name out loud: compliance mode and Vault Lock both remove your own ability to fix mistakes. That's the point, but it means the retention values have to be right first time, and I'd want governance mode in a staging environment to validate the settings before anything goes irreversible.

For the evidence the regulator wants: Backup Audit Manager for daily compliance reports against defined controls, CloudTrail for key usage, and the KMS rotation events in CloudTrail and EventBridge. And Macie on the buckets, so I can answer "where is the sensitive data" rather than assuming.

What I'd monitor: replication metrics if there's cross-Region replication with a recovery point objective attached, Backup Audit Manager findings, and GuardDuty — specifically for credential compromise, since that's the attack this design is built against.


Debug

"A Lambda function suddenly can't read from an S3 bucket. It worked yesterday. Walk me through your checks, in order."

  1. Read the actual error. Access denied from S3 and access denied from KMS are different problems with different fixes, and the message usually says which.
  2. CloudTrail for the failing call — look at the errorCode and the principal. Confirm it's the role I think it is.
  3. Did anything change? CloudTrail for recent PutBucketPolicy, PutRolePolicy, AttachRolePolicy, PutKeyPolicy, and any SCP changes in Organizations. "It worked yesterday" means something changed, and finding the change beats reasoning from first principles.
  4. Walk the evaluation chain in order: is there an explicit deny anywhere — identity policy, bucket policy, SCP, RCP, permissions boundary, session policy? Deny beats everything, so it's the cheapest thing to rule out.
  5. Then the ceilings: does the SCP allow this service at all? Is there a permissions boundary on the role?
  6. If the object is encrypted with a KMS key, check the key policy, not just the IAM policy — because an IAM allow means nothing if the key policy doesn't enable IAM policies.
  7. If it's a VPC-attached Lambda, check the network path too: endpoint or NAT, the endpoint policy, the route table, the security group.

I'd use IAM Policy Simulator or Access Analyzer somewhere around step 4 to check my reasoning rather than trusting it.

"Objects are disappearing from a bucket that compliance says is immutable. Both things appear to be true. Explain."

They probably both are true, and there are two candidates.

The first is a delete marker. On an Object Lock–protected object, a simple delete with no version ID returns 200 OK and inserts a delete marker, which becomes the current version. The object vanishes from a normal listing and the locked version is still there, still protected. So "disappeared" and "cannot be deleted" are both correct. I'd confirm with list-object-versions.

The second is a lifecycle rule, and that one is real deletion. A bucket policy can't stop a lifecycle rule — even a deny-everything policy doesn't touch it. So I'd check the lifecycle configuration for an expiration action, including noncurrent-version expiration, which people forget they wrote.

The distinguishing evidence is list-object-versions: if the versions are there under a delete marker, it's cosmetic. If they're gone, it's lifecycle, and I'd go to CloudTrail and S3 server access logs to confirm.

"A dashboard behind CloudFront is returning TLS errors after a Region migration to eu-west-2."

First thing I'd check is where the ACM certificate lives. Certificates are Regional resources and can't be copied between Regions, and CloudFront specifically requires the certificate in us-east-1. If the migration moved everything to eu-west-2 including the certificate, that's the whole bug.

You'd need two certificates: one in us-east-1 for CloudFront, one in eu-west-2 for the load balancer. That trips people up because it looks like an exception to "keep everything in one Region", and it is.


Red flags

Answers that sound confident and lose the offer.

"We use IAM policies for everything, so we don't need resource policies." Misses that cross-account access requires both sides, and that KMS actively requires the resource policy — an IAM allow alone does nothing on a key. It also gives up the union, which is often the simplest way to grant same-account access.

"An SCP gives the sandbox accounts EC2 access." An SCP never grants anything. It's a ceiling. Saying it this way signals you haven't internalised the model, and the interviewer will probe until it shows.

"We rotate keys, so we're covered." Rotation doesn't re-encrypt and doesn't help with a compromised data key. Saying this in a security review is how a real gap survives an audit.

"We denied delete in the bucket policy, so the data is immutable." Permissions aren't immutability. The lifecycle carve-out is documented and specific. The correct answer names Object Lock.

"Shield Standard is enabled on all our critical resources." There's nothing to enable — it's automatically included at no extra cost. It's a small thing, but it tells an interviewer you learned the service from a slide rather than the documentation.

"We put the WAF on the Network Load Balancer." You can't. If you say it without hesitating, the interviewer now doubts everything else you said about the perimeter.

"We keep the security tooling in the management account so it's centralised." Sounds tidy, but SCPs don't apply there, so you've put your most privileged workloads outside your own guardrails. The right answer is a delegated administrator account.

"It depends" — with nothing after it. Which brings us to the last one.


The one where "it depends" is correct

"Should we use S3 Object Lock in compliance mode?"

It depends, and here's what I'd ask you.

What exactly is the retention obligation, in writing? Compliance mode can't be shortened and can't be overridden by anyone including root — the documented way to delete early is to close the AWS account. So I need the number to be right the first time, from the regulation, not from an estimate.

Second, do we know the retain-until date at write time? If the clock starts on a business event — contract completion, claim resolution — then variable retention with an event hold is the better fit than committing to a date we'll have to work around.

Third, is anyone going to need to delete this for a legitimate reason? A data subject erasure request under a privacy regime, for instance, can collide directly with a financial retention obligation. That's a legal conflict, not an architecture one, and I'd want it resolved before I make something irreversible.

And fourth, have we tested the settings in governance mode first? Because governance mode is designed to be reversible and compliance mode isn't, so validating in governance costs nothing and getting compliance wrong costs seven years.

If the answers are "the regulation says seven years, the clock starts at write, nothing may be deleted, and yes we've tested it" — then compliance mode, absolutely. But knowing that this question is under-specified is more important than knowing what compliance mode does.