Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Starts this lesson and continues through 18 more to the end of certification prep.
Task statement 2.1 lists "Horizontal scaling and vertical scaling" and "Design principles for microservices (for example, stateless workloads compared with stateful workloads)" as separate knowledge items. They are not separate. The second one decides the first.
The course intro video for this exam calls out horizontal vs vertical as the thing that "usually confuses most students". It shouldn't, once you see that the interesting question isn't "which is better" but "which are you allowed to use".
| Vertical | Horizontal | |
|---|---|---|
| What changes | the instance gets bigger | there are more instances |
| Ceiling | the largest instance type available | effectively none |
| Downtime to apply | usually a stop/start | none — add capacity alongside |
| Resilience gain | none — still one instance | the point of the exercise |
| Precondition | none | the workload must be stateless |
| Exam signal | "the database is CPU-bound", legacy single-node software | "handle unpredictable traffic", "no single point of failure" |
Two consequences worth stating plainly:
⚠️ Vertical scaling buys performance, never availability. A bigger instance is still one instance in one Availability Zone. If a stem asks for both "handle more load" and "survive an AZ failure", vertical scaling answers only half, and the half it answers is the easy one.
⚠️ Horizontal scaling requires statelessness, and that is the actual work. Adding instances is trivial. Making the application tolerate any request landing on any instance is the engineering.
If a request can land on any instance, session state cannot live in instance memory. The standard destinations, and the exam expects you to name one:
| State | Where it goes |
|---|---|
| User session | ElastiCache, DynamoDB, or a signed cookie |
| Uploaded files | S3 (not the instance's local disk) |
| Shared file access across instances | EFS |
| Application data | RDS / Aurora / DynamoDB |
| Anything you can lose | instance store — and only then |
⚠️ Sticky sessions are a workaround, not a design. Session affinity on the load balancer keeps a user pinned to one instance so in-memory session state keeps working. It also means losing that instance loses that user's session, and it unbalances the fleet. On the exam, if an option offers sticky sessions and another option offers externalising session state, the second is the better architecture. Sticky sessions are the right answer only when the stem says the application cannot be modified.
That last case is real and the exam guide anticipates it — task statement 2.2 asks for skills in "Using AWS services that improve the reliability of legacy applications and applications not built for the cloud (for example, when application changes are not possible)".
From What is Amazon EC2 Auto Scaling?, verified 2026-09-22:
"Amazon EC2 Auto Scaling helps you ensure that you have the correct number of Amazon EC2 instances available to handle the load for your application. You create collections of EC2 instances, called Auto Scaling groups."
The three numbers that define a group, verbatim:
"You can specify the minimum number of instances in each Auto Scaling group, and Amazon EC2 Auto Scaling ensures that your group never goes below this size. You can specify the maximum number of instances … ensures that your group never goes above this size. If you specify the desired capacity … Amazon EC2 Auto Scaling ensures that your group has this many instances."
Minimum, maximum, desired. Scaling policies move desired between min and max.
⚠️ The most common real-world failure is max set too low. The scaling policy fires, the group is
already at maximum, and nothing happens. It looks like Auto Scaling is broken. It isn't — it's doing
exactly what you configured. Same shape as min set to 1 and wondering why there's no redundancy.
Verbatim, because each answers a different exam question:
Health checks and replacement — "Amazon EC2 Auto Scaling automatically monitors the health and availability of your instances using EC2 health checks and replaces terminated or impaired instances to maintain your desired capacity." Plus "you can define custom health checks that are specific to your application to verify that it's responding as expected. If an instance fails your custom health check, it's automatically replaced."
This is the self-healing answer. If a stem says "unhealthy instances must be replaced automatically", that's an Auto Scaling group — even at a fixed size. A group with min = max = desired = 3 does no scaling at all and is still worth having, because it replaces failures.
Balancing across AZs — "You can specify multiple Availability Zones for your Auto Scaling group, and Amazon EC2 Auto Scaling balances your instances evenly across the Availability Zones as the group scales. This provides high availability and resiliency by protecting your applications from failures in a single location."
Recall SAA0 lesson 4: "A subnet must reside within a single Availability Zone." So "multiple
Availability Zones" means multiple subnets on the group. Three AZs, three subnets.
Mixed instance types and purchase options — "Within a single Auto Scaling group, you can launch multiple instance types and purchase options (Spot and On-Demand Instances), allowing you to optimize costs through Spot Instance usage." This is a Domain 4 answer living in a Domain 2 service: a baseline of On-Demand plus Spot for the elastic portion.
Spot replacement — "If your group includes Spot Instances, Amazon EC2 Auto Scaling can automatically request replacement Spot capacity if your Spot Instances are interrupted. Through Capacity Rebalancing, Amazon EC2 Auto Scaling can also monitor and proactively replace your Spot Instances that are at an elevated risk of interruption."
Load balancer integration — "Whenever instances are launched or terminated, Amazon EC2 Auto Scaling automatically registers and deregisters the instances from the load balancer." You do not wire this up yourself, and an option that says you must is wrong.
Instance refresh — "provides a mechanism to update instances in a rolling fashion when you update your AMI or launch template. You can also use a phased approach, known as a canary deployment, to test a new AMI or launch template on a small set of instances before rolling it out." This is the immutable infrastructure answer that task statement 2.2 asks about: you don't patch instances, you replace them.
Lifecycle hooks — "useful for defining custom actions that are invoked as new instances launch or before instances are terminated." And for stateful workloads: "Lifecycle hooks also offer a mechanism for persisting state on shut down. … you can also use scale-in protection or custom termination policies to prevent instances with long-running processes from terminating early."
⚠️ Scale-in protection is the answer to "our batch jobs get killed mid-run when the fleet scales in." Not "turn off scaling".
Cost: "There are no additional fees with Amazon EC2 Auto Scaling … You only pay for the AWS resources … that you use."
⚠️ I did not retrieve the scaling-policy comparison (target tracking, step scaling, simple
scaling, scheduled, predictive) from the page I fetched — it lists features, not policy types. Read
docs.aws.amazon.com/autoscaling/ec2/userguide/scale-your-group.html before relying on the details.
What the exam guide itself gives you, under 2.1 skills, is "Determining scaling strategies for components used in an architecture design", and under 3.2 skills, "Identifying metrics and conditions to perform scaling actions". So the examinable shape is: what metric, and what threshold? A known daily pattern points at scheduled scaling; an unknown one points at tracking a metric.
The Well-Architected general design principle (General design principles, verified 2026-09-21) is the reasoning the exam rewards:
"Stop guessing your capacity needs: If you make a poor capacity decision when deploying a workload, you might end up sitting on expensive idle resources or dealing with the performance implications of limited capacity. With cloud computing, these problems can go away. You can use as much or as little capacity as you need, and scale in and out automatically."
So the scaling ladder, in order of decreasing operational overhead:
⚠️ Auto Scaling is a control-plane activity, and the DR whitepaper says so explicitly in the context of failover: "Because Auto Scaling is a control plane activity, taking a dependency on it will lower the resiliency of your overall recovery strategy. It is a trade-off." (Disaster recovery options in the cloud, verified 2026-09-22)
That's lesson 6's territory, but note it now: the thing that makes your steady state cheap is the thing you shouldn't lean on during a disaster. AWS names the alternative — "provision sufficient capacity such that the recovery Region can handle the full production load as deployed. This statically stable configuration is called hot standby."
Static stability is the idea: a system that doesn't need to do anything — call an API, scale out, provision — in order to survive the failure.
max.