AWS Training
Modules Listen All tracks

← Design Resilient Architectures

Starts this lesson and continues through 19 more to the end of certification prep.

Decoupling — queues, topics, events and workflows

Why this lesson comes first

Task statement 2.1 lists, under "Knowledge of": "Event-driven architectures", "Queuing and messaging concepts (for example, publish/subscribe)", and "Workflow orchestration (for example, AWS Step Functions)". Under "Skills in": "Determining the AWS services required to achieve loose coupling based on requirements".

Four services do this job and they are not interchangeable. The exam knows people reach for SQS reflexively, so it builds questions where SQS is plausible and wrong.

The distinguishing question is: how many consumers, and do they need to coordinate?

   one consumer, work to do          →  SQS          (queue: one message, one worker)
   many consumers, same message      →  SNS          (topic: fan-out, push)
   many consumers, filtered, routed  →  EventBridge  (bus: rules and patterns)
   steps that must happen in order   →  Step Functions (state machine)

SQS — a queue, and it is pull

From the Amazon SQS FAQs (verified 2026-09-21, and covered in detail in SAA0 lesson 5):

"Amazon SQS is a message queue service used by distributed applications to exchange messages through a polling model, and can be used to decouple sending and receiving components."

The two words that decide questions: polling (consumers pull; SQS never pushes to you) and decouple.

The numbers you need, all verbatim from the FAQ:

Standard FIFO
Ordering "loose-FIFO … attempts to preserve the order" "the exact order in which messages are sent and received"
Delivery at-least-once → consumers must be idempotent "exactly-once processing"
Throughput effectively unlimited 3,000/s batched · 300/s unbatched; high-throughput mode 70,000/s unbatched
Setting Value
Retention 1 minute – 14 days, default 4 days
Max message 1 MiB text; Extended Client Library references payloads "as large as 2 GiB"

⚠️ The claim-check pattern. A queue is not a file transport. For anything large, put the object in S3 and the reference in the message — which is exactly what the Extended Client Library does. Any exam option that pushes a 50 MB payload through SQS is wrong on the 1 MiB limit alone.

Visibility timeout is the mechanism behind at-least-once: "a period of time during which Amazon SQS prevents other consuming components from receiving and processing a message." Set it shorter than your processing time and you get duplicate processing — covered in SAA0 lesson 5 and drilled in that module's lab.

Long polling "doesn't return a response until a message arrives in the message queue, or the long poll times out", at "no additional charges compared to short polling". It is the answer to "reduce cost and empty receives", every time.

Dead-letter queues catch messages the consumer "is unable to consume … successfully". Without one, a poison message retries until retention expires — up to 14 days of a failing job.

SNS — a topic, and it is push

From What is Amazon SNS?, verified 2026-09-22:

"Amazon Simple Notification Service (Amazon SNS) is a fully managed service that provides message delivery from publishers (producers) to subscribers (consumers). Publishers communicate asynchronously with subscribers by sending messages to a topic, which is a logical access point and communication channel."

Subscriber endpoint types, verbatim — this list is examinable because it defines what SNS can and cannot reach:

And SNS spans both "Application-to-Application (A2A) and Application-to-Person (A2P) messaging" — which is why it's the answer for "alert the on-call engineer" as well as "notify three services".

Fan-out — the pattern the exam tests

"The Fanout scenario is when a message published to an SNS topic is replicated and pushed to multiple endpoints, such as Firehose delivery streams, Amazon SQS queues, HTTP(S) endpoints, and Lambda functions. This allows for parallel asynchronous processing."

AWS's own worked example is worth knowing because it's the canonical exam architecture:

"you can develop an application that publishes a message to an SNS topic whenever an order is placed for a product. Then, SQS queues that are subscribed to the SNS topic receive identical notifications for the new order. An … EC2 server instance attached to one of the SQS queues can handle the processing or fulfillment of the order. And you can attach another … EC2 server instance to a data warehouse for analysis of all orders received."

⚠️ SNS + SQS together is the answer more often than either alone. SNS gives you fan-out; the SQS queue in front of each consumer gives you durability and retry. A bare SNS→Lambda subscription drops the message if the Lambda is failing and retries are exhausted; SNS→SQS→Lambda does not, because the message sits in the queue. If a stem says "multiple systems must each process every event, and none may be lost", that's fan-out into queues.

AWS also names a second use: replicating production traffic into a test environment by subscribing an extra queue — with the caveat "Make sure you consider data privacy and security before you send any production data to your test environment."

SQS vs SNS — the one-line discriminator

SQS SNS
Model poll (consumer pulls) push (SNS delivers)
Consumers per message one (whichever worker gets it) many (every subscriber)
Storage messages persist up to 14 days no queue; delivery is attempted
Answers "work to be done" "something happened"

If the requirement says every consumer must get every message → SNS. If it says the work must be done once, by whichever worker is free → SQS.

EventBridge — routing on content

The exam guide lists event-driven architectures under 2.1 but names EventBridge only indirectly.

⚠️ I did not fetch the EventBridge documentation for this lesson, so this section makes no claim about its quotas, the schema registry, archive/replay, or the difference between the default bus, custom buses and partner buses. Read docs.aws.amazon.com/eventbridge/latest/userguide/eb-what-is.html before relying on specifics.

What is safe to reason about, and what the exam actually tests, is the shape: EventBridge is a bus with rules that match on event content and route to targets. So the discriminator against SNS is filtering and routing logic. SNS fans one message out to all subscribers; EventBridge decides, per event, which targets should see it. If the stem describes routing different event types to different consumers without the producer knowing about them, that's a bus, not a topic.

Step Functions — when order matters

From What is Step Functions?, verified 2026-09-22:

"With AWS Step Functions, you can create workflows, also called State machines, to build distributed applications, automate processes, orchestrate microservices, and create data and machine learning pipelines."

"Step Functions is based on state machines and tasks. … state machines are called workflows, which are a series of event-driven steps. Each step in a workflow is called a state. … Instances of running workflows performing tasks are called executions."

Standard vs Express — memorise this table

Verbatim, and it is the highest-yield Step Functions fact:

Standard Express
Execution semantics exactly-once at-least-once
Maximum duration one year five minutes
Execution rate 2,000 per second 100,000 per second
State transition rate 4,000 per second "Nearly unlimited"
Pricing "Priced by state transition" "Priced by number and duration of executions"
History "See execution history in Step Functions" "Send execution history to CloudWatch"
Integration patterns Request Response, Run a Job (.sync), Wait for Callback Request Response only

AWS's framing: Standard is "ideal for long-running, auditable workflows"; Express is "ideal for high-event-rate workloads, such as streaming data processing and IoT data ingestion."

⚠️ Express workflows cannot wait for a callback. "Express Workflows only support Request Response integrations." So any requirement involving a human approval step, or waiting for an external job to finish, forces Standard. That single constraint answers a lot of questions.

The four exam-shaped capabilities

From the use-case list, verbatim where quoted:

  1. Error handling — "You can retry failed tasks, or catch failed tasks and automatically run alternative steps" (Retry / Catch). This is the argument for Step Functions over chaining Lambdas by hand: the retry logic is declarative, not code you have to write and test.
  2. Human in the loop — "Step Functions can include human approval steps in the workflow", using "a callback and a task token". Standard only.
  3. Parallel — a Parallel state "can process input data in parallel steps".
  4. Map — "Using a Map state, Step Functions can run a set of workflow steps on each item in a dataset. The iterations run in parallel."

And the integration surface is enormous: "Over two hundred services" via AWS SDK integrations, plus optimized integrations for API Gateway, Athena, Batch, Bedrock, CodeBuild, DynamoDB, ECS/Fargate, EKS, EMR, EventBridge, Glue, Lambda, MediaConvert, SageMaker AI, SNS, SQS and Step Functions itself.

Choosing, under exam conditions

The stem says Answer
"decouple the front end from the workers so spikes don't drop requests" SQS
"each of three downstream systems must receive every event" SNS fan-out, ideally into SQS queues
"notify the operations team by SMS and email when the alarm fires" SNS (A2P)
"route different event types to different targets based on content" EventBridge
"these five steps must run in order, with retries, and one needs manager approval" Step Functions Standard
"process 50,000 IoT events per second through a short workflow" Step Functions Express
"strict ordering and no duplicates" SQS FIFO — check the throughput ceiling first
"a malformed message is blocking the queue" dead-letter queue

The deeper point: coupling is a failure domain

Task statement 2.1's skills list includes "Designing event-driven, microservice, and/or multi-tier architectures based on requirements". The reason the exam cares is not fashion. It's that a synchronous call couples the caller's availability to the callee's.

If your web tier calls the payment service synchronously and the payment service slows down, your web tier's threads block, its connection pool exhausts, and your homepage goes down because payments are slow. A queue converts that from an outage into a backlog.

Design rule, and it's worth being able to say out loud: a queue doesn't make anything faster. It changes the failure mode from rejection to delay — and delay is almost always the cheaper failure. The cost is that your system is now eventually consistent from the user's point of view, and the UI has to stop promising synchronous completion.

Check yourself

  1. Three services must each process every order event, and no event may be lost if one service is down. What's the architecture?
  2. A workflow needs a manager to approve a refund before continuing. Standard or Express?
  3. A stem requires strict ordering at 10,000 messages per second. What do you check?
  4. Your web tier's threads block whenever the payment service is slow. Name the fix and the trade-off.
  5. What's the difference between SNS and EventBridge in one sentence?
Answers
  1. SNS fan-out into three SQS queues, one per consumer. SNS gives every subscriber a copy; the queue in front of each consumer gives durability while that consumer is down. SNS alone would attempt delivery and eventually give up.
  2. Standard. Human approval uses the "Wait for Callback" pattern with a task token, and "Express Workflows only support Request Response integrations." Express also maxes out at five minutes, which no approval step respects.
  3. Whether FIFO high-throughput mode is available. Default FIFO is "up to 3,000 messages per second with batching or up to 300 … without"; high-throughput mode is "up to 70,000 … without batching." If only default FIFO is offered, 10,000/s is not achievable.
  4. Put a queue between them. The trade-off: the interaction becomes asynchronous, so the system is eventually consistent and the UI can no longer promise the payment completed synchronously.
  5. SNS pushes one message to every subscriber; EventBridge routes each event to the targets whose rules match its content.

Teaching this section

Next →Scaling — horizontal, vertical, and why state is the constraint