Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
The quiz tests recall. This rehearses speech. Read the answers out loud. Nothing here should take more than 90 seconds to say.
Resilience is the area where interviewers probe hardest, because it's where hand-waving is easiest to spot. The tell they're listening for is whether you name a trade-off or just a service.
"What's the difference between high availability and fault tolerance?"
Highly available means it recovers quickly — there's a gap, but it's short. Fault tolerant means it keeps running with no interruption at all, which costs more because you're paying for redundancy that's idle most of the time.
RDS Multi-AZ is the clean example. It's highly available: failover is automatic, but there is a gap while the standby is promoted, and open connections drop. It isn't strictly fault tolerant. If someone tells me they need zero dropped connections, a standby that has to be promoted doesn't meet that.
"Horizontal or vertical scaling?"
Different jobs. Vertical makes the instance bigger — it buys performance, it has a ceiling at the largest instance type, and it usually needs a stop-start. Horizontal adds instances — effectively no ceiling, no downtime to add capacity.
The thing I'd emphasise is that vertical scaling buys you no availability at all. A bigger instance is still one instance in one Availability Zone.
And horizontal has a precondition: the workload has to be stateless. Adding instances is trivial; making the app tolerate any request landing on any instance is the actual work.
"Where does session state go, then?"
Out of the instance. ElastiCache or DynamoDB for sessions, S3 for uploads, EFS if several instances genuinely need the same filesystem, and the database for application data.
Sticky sessions are the workaround, not the design — they mean losing an instance loses those users' sessions, and they unbalance the fleet. I'd only reach for them when the application can't be modified, which does happen with legacy systems.
"SQS or SNS?"
SQS is a queue and it's pull — one message goes to one worker, whoever's free. SNS is a topic and it's push — every subscriber gets a copy.
So: "there is work to be done" is SQS. "Something has happened" is SNS.
And in practice the answer is often both. SNS fans out to multiple SQS queues, one per consumer, so every system gets the event and each one has a durable buffer if it's down.
"What does Multi-AZ give you that read replicas don't, and vice versa?"
Multi-AZ is a synchronous standby in another AZ for availability and durability, with automatic failover. Read replicas are asynchronous copies for read scaling, and promotion is manual.
The thing people get wrong: the Multi-AZ standby cannot serve reads. It's not load-sharing. So if a database is both a single point of failure and overwhelmed by reads, that's two problems and you need both features.
"Walk me through how you'd find single points of failure in an architecture."
I'd walk the request path and at every hop ask what happens if that one thing dies.
Client, DNS, load balancer, compute, database, downstream dependencies. At each one: single instance means an Auto Scaling group across at least two AZs. Single AZ means subnets in two or more AZs — everywhere, because a subnet is zonal. Single-AZ RDS means Multi-AZ. One endpoint with no health check means Route 53 health checks and failover.
Then the two that aren't on the diagram. Connection limits — a serverless front end can exhaust a database's connections without any component actually failing. And service quotas, especially in a DR Region, which is a single point of failure you can't see until you try to scale into it.
The follow-up they ask next: "Which of those have you actually hit?" Be honest here. If you've hit the connection one, describe it — it's the most common and the most credible.
"Why isn't an Auto Scaling group enough for a Regional outage?"
Because it doesn't span Regions. Neither does Multi-AZ, neither does an Auto Scaling group's AZ balancing. Everything inside a Region protects you from losing an instance or a zone, and nothing more.
The moment the requirement says "survive losing a Region", you're choosing between backup-and-restore, pilot light, warm standby, and multi-site active-active — and which one depends entirely on the RTO and RPO.
The follow-up they ask next: "So what does an Auto Scaling group give you in a DR Region?" Careful answer: it gives you the ability to scale up, but AWS explicitly warns that Auto Scaling is a control plane activity, and depending on it lowers the resiliency of your recovery strategy. If you want to be safe, you provision full capacity in advance — that's hot standby, and the property is static stability.
"Explain the difference between pilot light and warm standby."
Both have your data replicating to the DR Region and a copy of your infrastructure there. The difference is whether it can serve a request right now.
Pilot light: the core infrastructure is on — databases, storage — but the application servers are switched off, or not even deployed. You have to turn things on before it can handle traffic.
Warm standby: everything is running, just scaled down. It can take traffic immediately, at reduced capacity, and you only need to scale up.
The one-line test I use is exactly that: can it serve a request right now without me doing anything?
The follow-up they ask next: "Which would you recommend?" It depends on the RTO. If the business can tolerate half an hour, pilot light is materially cheaper. If they need minutes, warm standby. And I'd push back on anyone who states an RTO without having tested a restore, because an untested RTO is a guess.
"What's the data plane rule and why does it matter?"
AWS divides services into a data plane, which delivers the actual service, and a control plane, which configures the environment. And the guidance is to use only data plane operations in your failover path, because data planes have higher availability design goals.
The reason it matters is timing. A Regional event is exactly when the control plane is under stress. So if your failover procedure depends on creating resources, changing configuration, or scaling out, you're depending on the least reliable part of the system at the worst possible moment.
Concretely: Route 53 health checks driving DNS failover is data plane — good. Application Recovery Controller routing controls are data plane — also good, and they're the right way to do a manually-triggered failover. Changing Route 53 weights is control plane. Global Accelerator traffic dials are control plane. Auto Scaling is control plane.
The follow-up they ask next: "So how do you do a manual failover properly?" Application Recovery Controller lets you create Route 53 health checks that don't actually check anything — they're on-off switches you control through a data plane API. So you get a deliberate human trigger with a highly-available mechanism underneath.
"Should DR failover be automatic?"
Usually not the trigger. AWS is cautious about this and I agree with the reasoning: recovery time and recovery point are always greater than zero, so if you fail over on a false alarm you incur those losses for nothing.
So: manual trigger, automated steps. A human decides, and then it's one button. You get the judgment of a person and the reliability of automation, and you don't lose an hour of data because a health check flapped.
The follow-up they ask next: "When would you automate the trigger?" When the cost of a false positive is genuinely low — an active-active setup where shifting traffic away from a Region is cheap and reversible. The more expensive the failover, the more I want a human in the loop.
"We replicate continuously to another Region. Are we protected?"
Against losing the Region, yes. Against a data disaster, no — and this is the distinction I'd want to draw clearly.
Replication is faithful. If someone runs a destructive command at nine o'clock, it's in your DR Region a second later. AWS says it directly: continuous replication may not protect against data corruption or malicious attack as well as point-in-time backups do.
So you need both. Replication for the Region failure, and point-in-time backups plus versioning for the data failure. And for anything regulated, Object Lock — because a deny in a bucket policy doesn't prevent a lifecycle rule from deleting the data, and permissions aren't immutability.
The follow-up they ask next: "What's your RPO for the data disaster, then?" Honest answer: it's bounded by your backup frequency and, more importantly, by detection time. The recovery point is always some moment before you discovered the problem, and people forget to count detection in the number.
"When would you choose EKS over ECS?"
When you actually need Kubernetes. EKS is certified Kubernetes-conformant, so applications deploy without refactoring and you keep community tooling — Helm charts, operators, whatever the team already has. And it runs on-premises through EKS Anywhere or Hybrid Nodes.
Everything else points at ECS. It's less to operate, and it doesn't have a per-cluster charge.
The thing I'd correct if someone said it: Fargate isn't an alternative to ECS or EKS. It's an alternative to EC2, underneath either of them. Orchestrator is one axis, capacity is the other.
The follow-up they ask next: "When wouldn't you use Fargate?" When you need something the task definition doesn't support on Fargate — GPU access, privileged mode, specific instance types or kernel tuning. Then you take EC2 capacity back, deliberately.
"Traffic spikes drop requests, the database falls over under reads, and the business now wants to survive losing a Region. Design it."
Three problems, and I want to keep them separate because they need different things.
Dropped requests under spikes is a coupling problem. The front end is doing synchronous work at whatever rate the internet sends it. So I'd put an SQS queue between accepting the request and doing the work — the front end validates and enqueues, which it can do fast, and a spike becomes a backlog instead of a failure. Workers scale on queue depth.
Because it's a standard queue, delivery is at-least-once, so consumers have to be idempotent. I'd key the work on something stable from the request rather than generating an ID consumer-side. And a dead-letter queue, because without one a poison message retries for up to fourteen days.
Read load is read replicas. Not Multi-AZ — the standby serves no reads. I'd point reporting and read-heavy paths at the replica endpoint and be explicit with the team that replication is asynchronous, so read-your-own-writes won't hold. That surfaces as a bug report eventually, so it needs handling in the application.
Regional survival is the expensive one, and I'd want the RTO and RPO before committing. If it's hours, backup and restore with infrastructure as code — and I'd stress that without IaC the redeploy time will blow the RTO. If it's tens of minutes, pilot light with Aurora global database, which replicates sub-second and promotes in under a minute. If it's minutes, warm standby.
Trade-offs I'd name out loud. The queue makes the system eventually consistent from the user's point of view, so the UI has to stop promising synchronous completion. Read replicas introduce lag. And the DR tier is a straight cost decision that the business, not engineering, should make — my job is to price each option honestly.
What I'd monitor: queue depth and message age — age is the one that tells you whether workers are keeping up; replica lag; a DLQ alarm; and X-Ray for latency across service boundaries, because per- service CloudWatch metrics can't tell you which call is slow.
And what I'd do before calling it done: test the failover. Not document it — test it. AWS's own design principle is "improve through game days", and an untested DR plan is a hypothesis.
"Requests are timing out intermittently. Only some users. Walk me through it."
The order matters — I'm going cheapest-to-check first, and each step either explains "only some users" or rules it out.
"An Auto Scaling group isn't launching instances. Five things you check."
Desired against max — if they're equal, it's working as configured and the fix is raising max. Scaling activities for the actual error. How many subnets are attached, and are they in different AZs. Whether the instance type is available in those AZs. And service quotas, especially if this is a Region you don't normally run in.
I'd also check whether the CloudWatch alarm fired at all — sometimes Auto Scaling is innocent and the alarm never triggered.
"Failover to the standby database completed successfully, but the application stayed down."
Almost certainly held connections. A Multi-AZ failover drops every open connection, and the application has to notice, reconnect and retry. If the connection pool doesn't handle that — or the app caches the resolved IP — it sits there failing against a dead endpoint.
Short term, recycle the application tier. The durable fix is RDS Proxy, which automatically connects to the standby while preserving application connections. That's precisely what it's for, and it's why it belongs in a resilience conversation rather than a performance one.
"We use Multi-AZ, so our reads scale." They don't. The standby serves no read requests. This is the single most common RDS misconception and interviewers use it deliberately.
"We have Auto Scaling, so we're covered for a Regional failure." Auto Scaling doesn't span Regions — and AWS explicitly warns against depending on it during a failover, because it's a control plane activity.
"Our DR plan is documented and signed off." Documented isn't tested. AWS's own principle is game days. An untested RTO is a guess, and it's usually optimistic.
"We replicate to another Region, so we're protected." Against Region loss, yes. Against deletion or corruption, no — replication copies the mistake perfectly.
"We'll put the WAF on the Network Load Balancer." You can't. WAF is layer 7; an NLB is layer 4. If you say it confidently, the interviewer now doubts everything else you said about the perimeter.
"We'll use Fargate instead of ECS." Category error. Fargate is how ECS runs, not an alternative to it.
"Adding a queue will make it faster." It won't. A queue converts rejection into delay. That's usually the better failure, but it's not speed, and describing it as speed suggests you haven't thought about the consistency consequences.
"We're active-active across two Regions, so we're fine." Only if either Region can carry the whole load. Two Regions each at 60% cannot absorb each other, and that's a capacity planning question people skip.
"Should we go multi-Region?"
It depends, and I'd want four things before answering.
What's the actual RTO and RPO, in numbers, signed off by someone who owns the revenue? Not "minimal downtime" — a number. Because that single pair of numbers selects the strategy, and everything else is argument.
What's our definition of a disaster? AWS draws this line usefully: for the loss of one data centre, a well-architected highly available single-Region workload may only need backup and restore. Multi-Region is for when your definition extends to losing a Region, or when a regulator requires it.
Have we tested a restore? If nobody has, the current RTO is unknown, and I'd rather spend the first month of budget finding out than spend a year's budget on active-active that also hasn't been tested.
Can each Region carry the full load? If we're going active-active at 60% utilisation each, we haven't built DR, we've built two ways to fail at once.
And I'd add one thing people skip: multi-Region raises the data consistency question, which is the genuinely hard part. Write global, write local, or write partitioned — and write local with DynamoDB global tables means last-writer-wins reconciliation, which some businesses simply cannot accept.
If the answers are "RTO 15 minutes, RPO under a minute, signed off, tested, and we've sized for full load" — then yes, warm standby at minimum, probably active-active. But knowing this question is under-specified is more useful than knowing what active-active costs.