AWS Training
Modules Listen All tracks
0:00 0:00

← Whitepapers and FAQs

Starts this lesson and continues through 26 more to the end of certification prep.

Data and messaging — RDS and SQS FAQs

Why these two, together

RDS and SQS are the last two documents on the reading list, and they pair better than they look. Both are answers to "this component is a single point of failure", and they answer it in opposite directions: RDS by replicating the thing, SQS by decoupling the things around it.

The exam tests both, and there is exactly one sentence in the RDS FAQ that decides more questions than any other on this page. We'll get to it second.

⚠️ Same sourcing caveat: my fetches returned substantial answers from both pages on 2026-09-21, not the complete documents. Gaps are flagged, not filled from memory.

Amazon RDS

What it is, verbatim:

"Amazon Relational Database Service (Amazon RDS) is a managed service that makes it easy to set up, operate, and scale a relational database in the cloud."

The engines, verbatim: "RDS for PostgreSQL, RDS for MySQL, RDS for MariaDB, RDS for SQL Server, RDS for Oracle, or RDS for Db2" — plus Amazon Aurora.

Seven names. Aurora is listed separately because it's AWS's own engine rather than a managed third-party one, and that distinction shows up in questions about performance and scaling.

Multi-AZ versus read replicas — and the sentence that decides everything

Multi-AZ, verbatim: RDS "automatically provisions and maintains a synchronous standby replica in a different Availability Zone" to provide enhanced durability and automatic failover capability.

Read replicas, verbatim: they "make it easier to take advantage of supported engines' built-in replication functionality to elastically scale out beyond the capacity constraints of a single DB instance for read-heavy database workloads."

And now the sentence:

⭐ "A Multi-AZ standby cannot serve read requests. Multi-AZ deployments are designed to provide enhanced database availability and durability, rather than read scaling benefits."

That is the single highest-value sentence on this page. It exists in the FAQ because people assume the opposite — you're paying for a second instance, so surely it should be doing something. It isn't. It sits there synchronously replicating and waiting.

Multi-AZ Read replica
Replication synchronous asynchronous (engine's built-in)
Purpose availability and durability read scaling
Serves reads No Yes
Failover automatic manual promotion
Same Region different AZ, same Region can be cross-Region

⚠️ The two-part question. A stem describing both "the database is a single point of failure" and "read queries are overwhelming it" needs both features — Multi-AZ for the first, read replicas for the second. An option offering only one is incomplete, and an option offering Multi-AZ as a fix for read load is simply wrong per the quote above.

⚠️ And the reverse trap: an option offering read replicas for high availability. They're asynchronous and need manual promotion, so they're not an availability mechanism. Multi-AZ is.

Automated backups

"By default, Amazon RDS enables automated backups of your DB instance with a 7-day retention period" and you can modify this "to any number from 0 … to the desired number of days, up to 35."

Value
Default retention 7 days
Range 0 to 35 days
Retention of 0 disables automated backups

⚠️ Zero is a legal value and it turns backups off. That's the trap: "set the retention period to 0 to save costs" is a real option and it's a disaster. And 35 days is the ceiling — if a requirement says "retain backups for 90 days", automated backups alone cannot do it. You need manual snapshots, or AWS Backup with a longer lifecycle (SAA1 lesson 6).

⚠️ I did not retrieve, from this page: the backup window and its I/O suspension behaviour on single-AZ instances, the difference between automated backups and manual snapshots on instance deletion, point-in-time recovery granularity, encryption of existing unencrypted instances, or the Multi-AZ DB cluster (three-instance) variant. All examinable. Read the rest of the RDS FAQ.

Amazon SQS

What it is, verbatim:

"Amazon SQS is a message queue service used by distributed applications to exchange messages through a polling model, and can be used to decouple sending and receiving components."

Two words to notice. Polling — consumers pull, SQS does not push (that's SNS or EventBridge). And decouple — which is why SQS turns up in resilience questions, not just messaging ones.

Standard versus FIFO

All verbatim:

Standard FIFO
Ordering "a loose-FIFO capability that attempts to preserve the order of messages" "preserve the exact order in which messages are sent and received"
Delivery "at-least-once delivery" "exactly-once processing, which means that each message is delivered once and remains available until a consumer processes it and deletes it"
Throughput (effectively unlimited) 3,000 msg/sec with batching; 300 msg/sec without. High throughput mode: "up to 70,000 messages per second without batching"

⚠️ "At-least-once" means duplicates. A standard queue can deliver the same message more than once, which is why consumers must be idempotent. If a stem says "processing a message twice would double-charge the customer", either make the consumer idempotent or use FIFO — and if the stem also says ordering matters, FIFO.

⚠️ The FIFO throughput numbers are the constraint that eliminates FIFO. A stem requiring 10,000 messages per second rules out default FIFO (3,000 with batching) unless high throughput mode is on offer. Standard queues are the answer when volume is the binding requirement — which is exactly the trade-off AWS built the two types around: ordering or throughput.

Queue configuration numbers

Setting Value
Message retention 1 minute to 14 days, default 4 days
Maximum message size 1 MiB of text data
Larger payloads Extended Client Library — references payloads "as large as 2 GiB"

⚠️ Retention default is 4 days, maximum is 14. A requirement to hold unprocessed messages for a month cannot be met by SQS retention.

⚠️ 1 MiB is the message limit, and the pattern for bigger payloads is the claim check: put the object in S3, put the reference in the message. That's what the Extended Client Library does, and it's the right answer whenever a stem mentions large files going through a queue.

The three behaviours that make queues work

Visibility timeout:

"a period of time during which Amazon SQS prevents other consuming components from receiving and processing a message"

This is the mechanism behind at-least-once delivery. A consumer receives a message, the message becomes invisible, and if the consumer doesn't delete it before the timeout expires it becomes visible again and somebody else gets it. So a visibility timeout shorter than your processing time causes duplicate processing — which is a real production bug and a good exam question.

Long polling:

"Long polling doesn't return a response until a message arrives in the message queue, or the long poll times out"

And, per the FAQ, "no additional charges compared to short polling".

⚠️ Long polling is nearly always the right answer. Short polling returns immediately even when the queue is empty, so you pay for empty receives and you burn requests. Long polling costs no more and reduces both. If an option says "reduce SQS costs and empty responses", it's long polling.

Dead-letter queues:

"an Amazon SQS queue to which a source queue can send messages if the source queue's consumer application is unable to consume the messages successfully"

⚠️ The DLQ is the answer to "a malformed message is blocking the queue" or "we keep reprocessing the same failing message forever". Without one, a poison message retries indefinitely until retention expires.

⚠️ I did not retrieve, from this page: the visibility timeout maximum (commonly cited as 12 hours), the default visibility timeout, the maximum long-poll wait time, delay queues, FIFO message group IDs and deduplication IDs, or SQS encryption details. All examinable. Read the rest of the SQS FAQ and docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/.

Pairing them: the decoupling answer

The reason both documents are on one reading list becomes clear in the exam's most common resilience pattern:

"During traffic spikes the application drops requests and the database is overwhelmed."

The framework-shaped answer uses both ideas:

  1. Decouple — put an SQS queue between the front end and the workers, so a spike becomes a backlog instead of a failure. "can be used to decouple sending and receiving components."
  2. Scale the reads — read replicas for read-heavy load.
  3. Protect availability — Multi-AZ for the failover requirement, separately from the read load.
  4. Make consumers idempotent — because a standard queue is at-least-once.

That's four facts from two FAQ pages producing one design. Which is the argument for reading them: each fact is small, and the combination is the exam.

Check yourself

  1. Read queries are overwhelming your RDS instance. You enable Multi-AZ. What happens to the read load?
  2. A requirement says database backups must be retained for 90 days. Can automated backups do it?
  3. A stem needs strict ordering at 10,000 messages per second. What do you need to check?
  4. A consumer takes 90 seconds to process a message and the visibility timeout is 30 seconds. What happens?
  5. Your SQS bill is full of empty receives. What's the fix, and what does it cost?
Answers
  1. Nothing. "A Multi-AZ standby cannot serve read requests. Multi-AZ deployments are designed to provide enhanced database availability and durability, rather than read scaling benefits." You needed read replicas.
  2. No. The automated backup retention range is "0 … up to 35" days. For 90 days use manual snapshots or AWS Backup with a longer retention lifecycle.
  3. Whether FIFO high throughput mode is available/offered. Default FIFO is "up to 3,000 messages per second with batching or up to 300 messages per second without"; high throughput mode is "up to 70,000 messages per second without batching." If only default FIFO is on offer, 10,000/sec is not achievable and the requirement needs revisiting.
  4. The message becomes visible again after 30 seconds and another consumer receives it — duplicate processing, while the first consumer is still working. Raise the visibility timeout above the processing time, and make the consumer idempotent.
  5. Long polling — and it's "no additional charges compared to short polling." Free fix.

Teaching this section

← PreviousNetworking — VPC and Route 53 FAQsFinished →Cheat sheet, lab & quiz