Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
The quiz tests recall. This rehearses speech. Read the answers out loud. Nothing here should take more than 90 seconds to say.
Cost questions in interviews are rarely about prices. They're about whether you know where the money goes and what condition each cheap option depends on. The tell they're listening for is whether you start from evidence — the bill — or from a favourite service.
"How do you find out what's driving an AWS bill?"
Start broad, then drill. Cost Explorer first, grouped by service and then by usage type, to see what grew and when. Then the Cost and Usage Report — now CUR 2.0 in Data Exports — for the line items: which resource, which hour, which usage type. Cost allocation tags tell me whose it is, but only if someone activated them; tagging alone isn't enough.
And I'd look at the usage types specifically, because data transfer and NAT processing hide there. A line called something-DataTransfer-Regional-Bytes is cross-AZ traffic, and people are routinely surprised by it.
"Savings Plans or Reserved Instances?"
Savings Plans, by default — and AWS says so on the Reserved Instances page itself. A Compute Savings Plan gives up to 66% and follows the workload across instance families, Regions, and onto Fargate and Lambda. An EC2 Instance Savings Plan gives up to 72% but locks you to one family in one Region.
The one case where I'd pick a Reserved Instance is when I need capacity guaranteed in a specific AZ. A zonal RI reserves capacity; a Savings Plan never does.
"When would you use Spot?"
Anything that can be interrupted and restarted — batch, CI, analytics, stateless workers, test environments. Up to 90% off On-Demand. The deal is a two-minute interruption notice, delivered through EventBridge and instance metadata, on a best-effort basis — so you design for interruption rather than relying on the warning.
Two things I'd flag. If the interruption behaviour is hibernate, there's no two-minute warning. And Spot doesn't stack with Savings Plans — it's not covered and doesn't count toward the commitment.
"What's the cheapest S3 storage class?"
Deep Archive, but that's the wrong question. The right one is: how often is it read, how fast must it come back, and how long will it live? Every class below Standard has a minimum — 30 days for the IA classes, 90 for Glacier Instant and Flexible, 180 for Deep Archive — and the IA and Glacier Instant classes bill anything under 128 KB as 128 KB. Put short-lived or tiny objects in a "cheap" class and the bill goes up.
"Our NAT gateway is one of our top five line items. What do you do?"
First, find out what's going through it. NAT gateways charge per GB processed on top of the hourly charge, so the question is which destinations. If a lot of it is S3 or DynamoDB in the same Region, I add gateway endpoints — AWS states there's no additional charge for them, and the route table sends that traffic to the endpoint instead of the NAT. That's often most of the bill gone.
For other AWS services, interface endpoints — which do cost money, so I'd compare against the NAT processing rather than assume.
Then placement. If instances in one AZ are using a NAT gateway in another, that's cross-AZ transfer as well. AWS's own advice is a NAT gateway per AZ, or keep the resources in the same AZ as the gateway.
The follow-up they ask next: "Why not just replace it with a NAT instance?" You can — it swaps the per-GB charge for an instance charge. But you now patch it, size it, and script its failover. AWS recommends the gateway for availability, bandwidth and admin effort. I'd only use an instance for a low-traffic dev environment.
"One NAT gateway or one per AZ?"
It depends on the environment, and I'd say that out loud. For dev: one, because it halves or thirds the hourly charges and losing outbound access during an AZ event is acceptable. For production: one per AZ, because AWS describes that as the zone-independent design, and because heavy traffic from other AZs through a single gateway pays cross-AZ transfer on top.
"How do you choose between DynamoDB on-demand and provisioned?"
On the shape of the traffic. AWS now describes on-demand as the default and recommended mode — pay per request, no capacity planning — and it's what I'd start with for anything new or spiky. Provisioned bills for the capacity you provision, not what you use, so it only wins when traffic is steady and forecastable enough to keep utilisation high. With auto scaling, AWS suggests a 70% target.
The follow-up they ask next: "How would you size provisioned capacity?" Capacity units. One RCU is one strongly consistent read per second up to 4 KB, or two eventually consistent. One WCU is one write per second up to 1 KB. Round up per item. Then I'd ask whether the reads really need strong consistency, because eventually consistent halves the read capacity.
"Is Lambda cheaper than EC2?"
For idle-heavy workloads, yes — you pay per request and per GB-second, so the idle time costs nothing. For something busy around the clock, not necessarily; you're paying for every GB-second of a workload that would sit happily on a right-sized instance with a Savings Plan. And a Compute Savings Plan covers Lambda and Fargate as well as EC2, so the commitment decision and the compute decision can be separated.
"How do you connect 40 VPCs cheaply?"
"Cheaply" and "40" pull in different directions. Peering is the lowest-cost option — no per-GB processing, free within an AZ — but it's point-to-point and not transitive, and AWS's guidance is that it suits fewer than about ten VPCs. Forty in a full mesh is 780 connections.
So I'd use Transit Gateway as the hub, and accept the per-attachment hourly charge and per-GB processing. Then, if two particular VPCs exchange a lot of data, I'd add a direct peering link between just those two. Hub for manageability, peering for the heavy pair.
The follow-up they ask next: "What does Transit Gateway cost?" Per attachment per hour plus per GB processed — on the pricing page I checked, in US East (Ohio), five cents an attachment-hour and two cents a gigabyte. I'd always name the Region, because it varies.
"Direct Connect or VPN?"
VPN is fast to set up, billed per connection-hour plus normal transfer, and each tunnel tops out at 1.25 Gbps. If I need more over VPN, it's multiple VPN connections on a Transit Gateway with ECMP and dynamic routing.
Direct Connect is a physical circuit: port-hours plus data out, and data in is free. The data-out rate is lower than internet egress, and the path is consistent because it bypasses ISPs. So for large, steady egress to on-premises, Direct Connect pays for itself. For modest or temporary needs, VPN.
"We're a startup. The bill doubled this quarter. Walk me through a cost review."
Evidence first, then levers, then guardrails.
Evidence. Cost Explorer by service and usage type to see what doubled. CUR in Athena to name the resources. Activate cost allocation tags if they aren't — you can't manage what you can't attribute.
Levers, in the order they usually pay off.
Idle and oversized first, because it's pure waste: Compute Optimizer recommendations, Trusted Advisor's
unattached Elastic IPs and idle load balancers, gp2 volumes that should be gp3 — 20% cheaper per GB
and performance independent of size.
Commitments second, but only for the steady baseline. A Compute Savings Plan sized to what we run every hour, not to the peak. Spot for the interruptible burst.
Storage third. Lifecycle rules for data with a known pattern, Intelligent-Tiering where the pattern is unknown — watching the 128 KB and minimum-duration traps.
Network fourth. Gateway endpoints for S3 and DynamoDB, CloudFront in front of anything serving the internet, and a look at cross-AZ traffic.
Guardrails. AWS Budgets with forecasted alerts to catch it next time before the invoice, and budget actions on sandbox accounts so a runaway experiment gets stopped rather than just reported.
Trade-offs I'd name: commitments are hard to undo — Savings Plans can't be changed and RIs can't be cancelled — so I'd commit conservatively and top up. And I wouldn't touch production resilience to save money without the business signing off on the risk.
"Our budget alert fired, but we still overspent by 30%. Why?"
Three candidates. Budgets data is updated up to three times a day, roughly eight to twelve hours apart, and AWS says you can exceed the threshold before the notification arrives. The alert may have been on actual spend when it should have been forecasted. And an alert on its own stops nothing — if the goal was to stop spend, it needed a budget action applying a deny policy.
"We moved logs to Standard-IA and the storage bill went up."
Almost certainly small objects or short lifetimes. Anything under 128 KB is billed as 128 KB, and anything deleted before 30 days is billed for 30. Log files are often both. The fix is to aggregate small files before tiering, or keep short-lived data in Standard. Lifecycle rules also skip objects under 128 KB by default now, for exactly this reason.
"We bought Reserved Instances, but our EC2 bill didn't drop."
The RI isn't matching usage. A Reserved Instance is a billing discount applied to instances that match its attributes — instance type, Region, tenancy, platform. Wrong family, wrong Region, a different OS, or a zonal RI in the wrong AZ, and nothing matches. I'd check the RI utilisation report — and set an RI utilisation budget so we hear about it next time.
"Spot plus a Savings Plan gives us both discounts." It doesn't. Spot isn't covered by Savings Plans and doesn't count toward the commitment.
"We tagged everything, so chargeback is done." Not until the tags are activated for cost allocation — by the management account — and it takes up to 24 hours to show.
"Glacier Instant Retrieval is our archive tier; we restore from it when needed." There's nothing to restore. It serves in milliseconds. Only Flexible Retrieval and Deep Archive need a restore.
"We'll use Requester Pays to share the dataset publicly." Requester Pays blocks anonymous access. It's for sharing with other AWS accounts.
"Transit Gateway is always better than peering." More manageable at scale, not cheaper. Peering has no per-GB processing charge.
"Serverless means we pay nothing when idle." Only if it actually scales to zero. Aurora Serverless reaches zero ACUs only on newer engine versions; older ones bottom out above zero.
"Automated backups cover our seven-year retention." They max out at 35 days. Long retention needs manual snapshots or AWS Backup.
"Should we buy a three-year commitment?"
It depends, and I'd want four things before answering.
How steady is the baseline? I'd look at the lowest hourly usage over the last few months, not the average. Commit to what we run every hour of every day.
How likely is the architecture to change? If we're moving to containers or Graviton, a Compute Savings Plan survives that; an EC2 Instance Savings Plan or a Standard RI may not. The extra discount on the locked options isn't worth much if the usage walks away from it.
Do we need capacity guarantees? If yes, zonal RIs or Capacity Reservations. If no, Savings Plans.
What's the cash position? All Upfront generally saves the most, but it's cash now. No Upfront keeps cash, and for RDS it's only offered on one-year terms.
And the thing I'd say every time: commitments can't be undone — Savings Plan terms can't be changed after purchase, and RIs can't be cancelled. So I'd start with a one-year plan sized below the baseline, watch the utilisation budget, and add to it. Under-committing costs a little; over-committing costs a lot.