AWS Training
Modules Listen All tracks
0:00 0:00

← Whitepapers and FAQs

SAA0 Interview — the facts, spoken

The quiz tests recall. This rehearses speech. Read the answers out loud. Nothing here should take more than 90 seconds to say.

These are questions the subject matter justifies for a cloud or data engineering role. The facts in this module are the ones interviewers use as a quick competence check, because they're short, they're checkable, and people bluff them.


Warm-up

"What's the Well-Architected Framework, in a sentence?"

It's AWS's set of foundational questions for evaluating an architecture, organised into six pillars. And the thing I'd flag is that it's explicitly about trade-offs — AWS's own framing is that a review is a constructive conversation about architectural decisions, not an audit. So it's not a checklist you pass. It's a way of being deliberate about which quality you're optimising for.

Six pillars: operational excellence, security, reliability, performance efficiency, cost optimisation, and sustainability.

"Can a subnet span two Availability Zones?"

No. A subnet lives in exactly one Availability Zone. The VPC is the Regional thing; the subnet is the zonal thing.

Which means multi-AZ isn't a setting you turn on — it's multiple subnets. An Auto Scaling group across three AZs needs three subnets, a load balancer needs a subnet in every zone it serves, and the one-NAT-gateway-per-AZ pattern exists because a NAT gateway lives in a subnet.

"How many usable IPs in a /28 subnet?"

Eleven. AWS reserves five addresses in every subnet — the first four and the last one — for the router, DNS, and broadcast. So it's not the fourteen you'd get on-premises, where you'd only lose the network and broadcast addresses.

"Multi-AZ or read replicas — what's the difference?"

Different problems. Multi-AZ is a synchronous standby in another Availability Zone for availability and durability, with automatic failover. Read replicas are asynchronous copies for read scaling, and promotion is manual.

And the thing people get wrong: a Multi-AZ standby can't serve reads. It's not a load-sharing setup. So if a database is both a single point of failure and overwhelmed by read queries, that's two problems and you need both features.

"Standard or FIFO queue?"

Ordering or throughput — you pick. FIFO gives exact ordering and exactly-once processing, but caps at 3,000 messages a second with batching, or 300 without, unless you turn on high throughput mode. Standard is effectively unlimited but it's loose ordering and at-least-once delivery, so consumers have to be idempotent.

Most of the time the honest answer is standard plus idempotent consumers, because strict ordering is assumed far more often than it's actually required.


Depth

"Why isn't VPC peering enough for a hub-and-spoke network?"

Because peering isn't transitive. If A peers with B and B peers with C, A still can't reach C — that's not discouraged, it's unsupported. So you can't use a hub VPC as a router.

Which means a full mesh, and the arithmetic kills you: n VPCs need n times n-minus-one over two connections. Ten VPCs is forty-five peering connections, and each one needs route table entries at both ends. Twenty VPCs is a hundred and ninety.

That's the whole case for Transit Gateway. One attachment per VPC, and transitive routing works.

The follow-up they ask next: "So is peering ever the right answer?" Yes — two or three VPCs that need to talk, especially cross-account or cross-Region, where you want an explicit, auditable point-to-point link and you don't want the cost of a transit gateway. It's a scaling problem, not a correctness problem.

"Latency-based routing or geolocation routing?"

Depends whether the requirement is performance or compliance, and that distinction matters more than it sounds.

Latency-based routing optimises — it sends each user wherever is fastest, which is usually but not necessarily their own region. Geolocation routes on where the request actually came from, at continent, country or state level.

So if the requirement is "German users must be served from Frankfurt for data residency reasons", latency-based routing is wrong even though it would probably do the right thing. "Probably" isn't a control. You need the policy that routes on the attribute the regulator cares about.

The follow-up they ask next: "What if you need both?" You'd use geolocation as the outer constraint and then latency or weighted routing within a region's set of endpoints. The compliance boundary goes on the outside, because that's the one that can't be violated.

"Why can't you use a CNAME for your root domain?"

It's a DNS protocol restriction, not an AWS one — a CNAME can't coexist with the other records that have to exist at a zone apex, like SOA and NS.

Route 53's answer is the alias record, which is an AWS-specific extension. It works at the apex, and it tracks the target resource's IP addresses automatically, so when your load balancer's addresses change you don't have to do anything.

And it's free — there's no query charge for alias records pointing at ELB, CloudFront, Beanstalk, API Gateway or VPC endpoints. So it wins on function and on cost.

"Is S3 eventually consistent?"

Not any more. S3 delivers strong read-after-write consistency automatically, with no change to performance or availability — after a successful write, the next read gets the latest version.

I'd flag this one specifically because a lot of material still teaches eventual consistency and the workaround patterns for it. If someone's architecture has a consistency workaround in it, it's probably load-bearing for nothing.

The follow-up they ask next: "Does that mean you can use S3 as a database?" No — consistency isn't the same as transactions, and there's still no atomic multi-object operation, no conditional writes across keys, and the request-rate and latency profile is object storage, not OLTP.

"When would you choose S3 One Zone-IA?"

When the data is re-creatable. It's designed for 99.5% availability rather than 99.99% because it's in a single Availability Zone, and that's the trade you're making for the lower price.

So: derived datasets, thumbnails, transcoded outputs, a secondary copy of something that's replicated elsewhere. Not the authoritative copy of anything, because losing the AZ loses the data.

And I'd be careful with the word durability here — the durability design is the same eleven nines. It's availability that differs. People use those two words interchangeably and they mean different things.

"An SQS consumer is processing the same message twice. Why?"

Two candidates, and I'd check the cheap one first.

Most likely the visibility timeout is shorter than the processing time. The consumer receives the message, it goes invisible, and if it's not deleted before the timeout expires it becomes visible again and a second consumer picks it up — while the first one is still working. Fix is to raise the timeout above the worst-case processing time, or extend it while processing.

The other possibility is that it's just a standard queue behaving as documented. Standard queues are at-least-once, so duplicates are expected, not a bug. In which case the fix isn't the queue, it's making the consumer idempotent.

The follow-up they ask next: "How would you tell which one it is?" Compare the processing duration against the visibility timeout, and check whether the duplicate arrives roughly one timeout after the original. If it does, it's the timeout. If duplicates are random and immediate, it's at-least-once delivery.


Design

"Traffic spikes are causing dropped requests, and the database falls over under read load. Design the fix."

Let me separate the problems first, because there are three and they need different things.

Dropped requests during spikes is a coupling problem — the front end is doing synchronous work at the rate the internet feels like sending it. So I'd put a queue between accepting the request and doing the work. SQS, standard queue, and the front end's job becomes "validate and enqueue", which it can do fast. A spike becomes a backlog instead of a failure, and the workers scale on queue depth.

Because it's a standard queue, at-least-once, the consumers have to be idempotent — so I'd key the work on something stable from the request rather than generating an ID on the consumer side.

The read load is a separate problem and the answer is read replicas, because that's what they're for. I'd point the reporting and read-heavy paths at the replica endpoint and leave writes on the primary, and I'd be explicit with the team that the replica is asynchronous, so there's replication lag and read-your-own-writes won't hold.

And then availability, which is the third thing and the one people conflate with the second. Multi-AZ, because the standby gives automatic failover — and specifically not because it helps with reads, since it can't serve them.

Trade-offs I'd name: the queue makes the system eventually consistent from the user's point of view, so the UI has to stop promising synchronous completion. And read replicas introduce lag that will surface as a bug report — "I saved it and it's not there" — so that needs handling in the application, not just in the infrastructure.

What I'd monitor: queue depth and message age, because age is the one that tells you whether workers are keeping up; replica lag; and a dead-letter queue with an alarm on it, because without one a poison message retries for up to fourteen days.


Debug

"An Auto Scaling group is configured for three Availability Zones but instances only launch in one. What do you check?"

  1. How many subnets are attached to the group. A subnet is in one AZ, so three AZs needs three subnets — this is the most likely answer and it's the fastest to check.
  2. Whether the other subnets actually exist in different AZs. It's easy to create three subnets and put two of them in the same zone.
  3. Available IP addresses in each subnet. A small subnet fills up fast, and remember AWS takes five addresses off the top, so a /28 only holds eleven instances.
  4. Instance type availability in each AZ — not every type is offered in every zone, and a launch template pinned to one type can silently constrain you to one AZ.
  5. Capacity errors in the scaling activity history, which will usually name the actual reason.

"Your RDS automated backups aren't going back as far as compliance needs. What happened?"

Almost certainly the retention period. The range is 0 to 35 days, and the default is 7 — so if nobody changed it, you have a week.

And if the requirement is longer than 35 days, automated backups cannot meet it at all, no matter what you set. That needs manual snapshots on a schedule, or AWS Backup with a longer retention lifecycle.

The thing I'd check for first, though, is whether the retention is set to 0, because that's a legal value and it disables automated backups entirely. Someone setting it to zero to save money is a real thing that happens, and it looks like "backups aren't going back far enough" right up until you realise there aren't any.


Red flags

"S3 is eventually consistent, so we added a retry loop." Out of date. Strong read-after-write consistency is automatic now. The retry loop is dead code defending against something that can't happen, and it suggests the architecture hasn't been revisited.

"We enabled Multi-AZ to spread the read load." It can't serve reads. This is the single most common RDS misconception and interviewers use it deliberately.

"A /28 gives us 14 addresses." Five reserved, not two. Eleven. It's a small thing, but it's the kind of detail that tells someone whether you've actually sized a subnet or only read about it.

"We'll peer the VPCs through the shared services VPC." Peering isn't transitive. Saying this confidently means you haven't built it.

"We used a CNAME for the apex domain." You can't. If you say you did, either you used an alias record and called it a CNAME, or it isn't working.

"Spot Instances, because the question said most cost-effective." Cost optimisation doesn't license breaking a functional requirement. Spot needs the workload to tolerate interruption, and saying otherwise suggests you'd optimise a payment processor into unreliability.

"There are five pillars." Six. Sustainability has been in the Framework for years. It's a date-check on your sources more than a knowledge gap, but that's what the interviewer learns from it.


The one where "it depends" is correct

"Should we use S3 Intelligent-Tiering?"

It depends on whether you actually know the access pattern, and I'd want to ask.

If you genuinely don't know, or it changes — user-generated content, data whose relevance decays unpredictably — then yes, Intelligent-Tiering is the right answer, and the Framework agrees with it: the first general design principle is literally stop guessing your capacity needs, and this is the storage version of that.

But if you do know the pattern — accessed heavily for thirty days, then almost never, then archived for compliance — then a specific storage class with a lifecycle rule is usually cheaper, because you're not paying Intelligent-Tiering's monitoring charge to discover something you already knew.

So the questions I'd ask are: do we have access-pattern data, from Storage Lens or access logs? What's the average object size, since monitoring charges are per object and lots of tiny objects change the maths? And are the objects large enough and long-lived enough to clear the minimum storage durations on the colder classes, because transitioning then deleting early can cost more than staying put?

Knowing that this question is under-specified is the point. "Use Intelligent-Tiering" as a reflex is the same mistake as guessing the pattern — you've just outsourced the guess.