Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
The quiz tests recall. This rehearses speech. Read the answers out loud. Nothing here should take more than 90 seconds to say.
Performance is where interviewers check whether you measure before you buy. The tell they're listening for is whether you name the bottleneck first, or reach straight for a bigger instance.
"Object, file or block — how do you choose?"
By how the data is accessed. If it's addressed by key over HTTP and read by lots of clients, that's object storage — S3. If several instances need to mount the same filesystem, that's file — EFS for Linux, FSx when it's Windows or HPC. If it's a disk for one instance, that's block — EBS, or instance store if it's scratch.
The one I'd flag straight away is Windows. AWS states plainly that EFS isn't supported on Windows instances, so a Windows file share is FSx for Windows File Server, with SMB and Active Directory.
"gp2 or gp3?"
gp3, almost always. gp3 gives a flat 3,000 IOPS and 125 MiB/s baseline included in the price, lets me buy more IOPS or throughput without growing the volume, doesn't burst, and AWS prices it about 20% lower per GiB than gp2.
gp2 ties IOPS to size — three per GiB — and small volumes burst on credits. That's the thing that bites: a job that's fast for half an hour and then slow is usually a gp2 volume that's spent its credits.
"What's instance store for?"
Scratch. Buffers, caches, temporary data, or data that's replicated across a fleet anyway. It's physically attached to the host, it's included in the instance price, and it's fast.
The catch is that it survives a reboot and nothing else. Stop, hibernate, terminate, change the instance type, or lose the disk — it's gone. So I'd never put anything there that has to survive a stop-start.
"How do you give a Lambda function more CPU?"
You raise its memory. There's no separate CPU setting; AWS allocates CPU in proportion to memory, and at 1,769 MB you get the equivalent of one vCPU.
What I'd add is that it isn't automatically more expensive. You pay for memory times duration, so if a CPU-bound function finishes proportionally faster, the bill can stay flat. I'd measure it — Power Tuning does exactly that — rather than guess. And there's a ceiling: if the code is single-threaded, a second vCPU's worth of memory won't help.
"CloudFront or Global Accelerator?"
CloudFront caches content at the edge — static assets, video, anything cacheable over HTTP. Global Accelerator doesn't cache anything; it gives you two static anycast IPs and carries the connection over AWS's network to the nearest healthy Regional endpoint.
So if the requirement is "partners need to allowlist fixed IPs", or it's UDP, or failover is being slowed by clients caching DNS — Global Accelerator. If it's "reduce load on the origin and speed up image delivery" — CloudFront.
"Walk me through scaling a worker fleet off an SQS queue."
The obvious metric — messages in the queue — is the wrong one, and AWS says so: the queue length doesn't change in proportion to the number of instances.
What you want is backlog per instance: visible messages divided by in-service instances. And the target comes from the business: acceptable latency divided by average processing time per message. If each message takes a tenth of a second and ten seconds is acceptable, the target is 100 messages per instance. Then target tracking does the rest — it scales out when the backlog per instance is above 100 and in when it's below.
Two extras. AWS now recommends metric math over publishing a custom metric. And for long-running work I'd use scale-in protection so the group doesn't terminate an instance mid-message.
"Read replicas, Aurora Replicas, ElastiCache, DAX — when is each the answer?"
All four are read-side fixes, so the first question is whether the problem is actually reads.
RDS read replicas: asynchronous copies for read-heavy workloads and reporting. I create them manually — RDS doesn't auto scale them — and the application has to tolerate lag.
Aurora Replicas: up to fifteen, sharing the cluster volume, so lag is usually well under 100 milliseconds, and they're also the failover targets. That's the big difference from RDS Multi-AZ, where the standby serves no reads.
ElastiCache: when the same data is read over and over and sub-millisecond matters. Lazy loading if I only want to cache what's read and can live with some staleness; write-through if reads must never be stale; a TTL on both.
DAX: only for DynamoDB, only for eventually consistent reads, and only when there's a high repeat-read ratio — AWS says it performs best above 90% hit rate. It's explicitly not for write-heavy or strongly-consistent workloads.
"How do you size a Kinesis stream?"
In provisioned mode, each shard takes 1 MB/s or 1,000 records a second in, and 2 MB/s out. AWS's formula is the larger of write KiB/s over 1,024 and read KiB/s over 2,048, rounded up — and read bandwidth is write bandwidth times the number of consumers sharing it.
If traffic is unpredictable, I'd use on-demand, which handles up to double the previous 30-day peak — with the caveat that doubling within fifteen minutes can throttle.
And the thing people miss: if one shard is hot while the rest idle, more shards won't help. It's the partition key. Same key, same shard.
"Direct Connect or VPN?"
VPN is IPsec over the internet, two tunnels per connection, quick to set up, 1.25 Gbps per standard tunnel. Direct Connect is a physical connection that bypasses the internet — dedicated at 1, 10, 100 or 400 Gbps, or hosted through a partner from 50 Mbps up.
The practical question is time. AWS says reviewing a dedicated request can take up to 72 business hours, and then there's a physical cross-connect to order. So if the need is this week, it's VPN now and Direct Connect when it lands.
The follow-up they ask next: "Is Direct Connect encrypted?" I'd check rather than answer from memory. The docs I've read confirm MACsec is available on dedicated connections; for anything beyond that I'd go to the Direct Connect security documentation before committing in a design.
"A global mobile game: players worldwide, UDP game traffic, a leaderboard, and match history analytics. Design the performance path."
Four flows, and they want different services.
Game traffic is UDP and latency-sensitive, and players should land on a nearby Region. That's Global Accelerator — static anycast IPs, traffic onto AWS's network at the nearest edge. If players have to reach a specific game server, custom routing accelerators map users to a specific destination.
Game servers that talk to each other: a cluster placement group for low node-to-node latency, on an instance type with enhanced networking. I'd accept that cluster placement is one AZ, and handle availability at the match level, not the instance level.
The leaderboard is a ranked read-mostly structure — ElastiCache for Valkey or Redis OSS, because sorted sets and failover are there and Memcached has neither. The durable record goes in DynamoDB; the cache is write-through so rankings are never stale.
Match history streams into Kinesis Data Streams if several consumers need it in real time, or straight into Firehose if it's only going to S3. Firehose converts JSON to Parquet on the way in using a Glue schema. Then Athena on partitioned Parquet, and dashboards in Quick Sight.
What I'd monitor: shard-level throttling for hot keys, cache hit rate, and data scanned per Athena query, because that's the bill.
Trade-offs I'd name: write-through means cache churn; cluster placement trades AZ resilience for latency; and Firehose buffers, so analytics is near-real-time, not instant.
"An RDS database has high write latency. Walk me through it."
DiskQueueDepth climbing and WriteIOPS near the provisioned number says yes."Lambda functions keep failing with 'too many connections' against RDS."
Each concurrent execution environment opens its own connection, so a burst of concurrency becomes a connection storm. RDS Proxy pools and reuses connections, queues or throttles what it can't serve, and usually needs no code change. Two limits I'd mention: it attaches to the writer only, not a read replica, and it has to be in the database's VPC.
"Athena queries got slower and more expensive over six months."
Almost always data layout. I'd check how many files are in each partition — lots of small files are slow, and at scale they hit S3's per-prefix request rate and return SlowDown. I'd check the format: CSV reads every column, Parquet reads only what the query needs. And I'd check the partitioning matches the queries — hourly partitions for daily queries create exactly the small-file problem. The fix is compact, convert to Parquet, and partition for the common query.
"We'll put EFS on the Windows servers." Not supported. AWS says so in one line. It's FSx for Windows File Server.
"We'll use Max I/O mode on EFS for performance." AWS now calls Max I/O a previous generation mode with higher per-operation latency and recommends General Purpose for all file systems.
"We'll add read replicas — the database is slow on writes." Replicas scale reads. They do nothing for write latency.
"We'll add DAX to speed up our ledger." If the ledger needs strongly consistent reads or is write-heavy, AWS says DAX isn't ideal on both counts.
"Cluster placement group for high availability." It's the opposite. Cluster is one AZ, packed close together, for latency.
"We'll give the Lambda more CPU." There's no CPU setting. It's memory — and saying "CPU" suggests you haven't configured one.
"We'll put CloudFront in front so partners can allowlist the IPs." CloudFront gives you a domain name, not two fixed IPs. That's Global Accelerator.
"We'll order Direct Connect for next week's go-live." Port provisioning alone can take up to 72 business hours, plus a physical cross-connect. VPN first.
"We'll ship it on Snowball." AWS's Snowball Edge guide now says it's no longer available to new customers. Say DataSync, or check current options — don't recommend a service from memory.
"Should we cache it?"
It depends, and I'd want four things before answering.
What's the read-to-write ratio, and how often is the same item read? A cache pays off on repeat reads. DAX's own guidance is that it performs best above 90% hit rate. If every read is for a different key, a cache is cost and complexity with no benefit.
How stale can the answer be? If "never", it's write-through with a TTL, and I accept churn. If seconds or minutes, lazy loading with a TTL is simpler and caches only what's actually read. If the business says "strongly consistent", DAX is out entirely.
Is the bottleneck actually the database? If the database is waiting on I/O, the answer might be Provisioned IOPS or a bigger instance class, not a cache. If it's connections, it's RDS Proxy. A cache in front of a connection problem just moves the problem.
Can we afford a cold cache? Lazy loading survives a node loss — it's slower until it refills — but if the database can't survive the miss storm after a restart, the cache has become a dependency, and now it needs its own availability design.
If the answers are "90% of reads hit the same few thousand items, seconds of staleness is fine, and the database is CPU-bound on those reads" — then yes, ElastiCache with lazy loading and a TTL. But knowing that "add a cache" is under-specified is worth more than knowing how to configure one.