AWS Training
Modules Listen All tracks

← Design Resilient Architectures

Starts this lesson and continues through 18 more to the end of certification prep.

Scaling — horizontal, vertical, and why state is the constraint

Why this lesson matters

Task statement 2.1 lists "Horizontal scaling and vertical scaling" and "Design principles for microservices (for example, stateless workloads compared with stateful workloads)" as separate knowledge items. They are not separate. The second one decides the first.

The course intro video for this exam calls out horizontal vs vertical as the thing that "usually confuses most students". It shouldn't, once you see that the interesting question isn't "which is better" but "which are you allowed to use".

The two directions

Vertical Horizontal
What changes the instance gets bigger there are more instances
Ceiling the largest instance type available effectively none
Downtime to apply usually a stop/start none — add capacity alongside
Resilience gain none — still one instance the point of the exercise
Precondition none the workload must be stateless
Exam signal "the database is CPU-bound", legacy single-node software "handle unpredictable traffic", "no single point of failure"

Two consequences worth stating plainly:

⚠️ Vertical scaling buys performance, never availability. A bigger instance is still one instance in one Availability Zone. If a stem asks for both "handle more load" and "survive an AZ failure", vertical scaling answers only half, and the half it answers is the easy one.

⚠️ Horizontal scaling requires statelessness, and that is the actual work. Adding instances is trivial. Making the application tolerate any request landing on any instance is the engineering.

Where the state goes

If a request can land on any instance, session state cannot live in instance memory. The standard destinations, and the exam expects you to name one:

State Where it goes
User session ElastiCache, DynamoDB, or a signed cookie
Uploaded files S3 (not the instance's local disk)
Shared file access across instances EFS
Application data RDS / Aurora / DynamoDB
Anything you can lose instance store — and only then

⚠️ Sticky sessions are a workaround, not a design. Session affinity on the load balancer keeps a user pinned to one instance so in-memory session state keeps working. It also means losing that instance loses that user's session, and it unbalances the fleet. On the exam, if an option offers sticky sessions and another option offers externalising session state, the second is the better architecture. Sticky sessions are the right answer only when the stem says the application cannot be modified.

That last case is real and the exam guide anticipates it — task statement 2.2 asks for skills in "Using AWS services that improve the reliability of legacy applications and applications not built for the cloud (for example, when application changes are not possible)".

EC2 Auto Scaling — what it actually guarantees

From What is Amazon EC2 Auto Scaling?, verified 2026-09-22:

"Amazon EC2 Auto Scaling helps you ensure that you have the correct number of Amazon EC2 instances available to handle the load for your application. You create collections of EC2 instances, called Auto Scaling groups."

The three numbers that define a group, verbatim:

"You can specify the minimum number of instances in each Auto Scaling group, and Amazon EC2 Auto Scaling ensures that your group never goes below this size. You can specify the maximum number of instances … ensures that your group never goes above this size. If you specify the desired capacity … Amazon EC2 Auto Scaling ensures that your group has this many instances."

Minimum, maximum, desired. Scaling policies move desired between min and max.

⚠️ The most common real-world failure is max set too low. The scaling policy fires, the group is already at maximum, and nothing happens. It looks like Auto Scaling is broken. It isn't — it's doing exactly what you configured. Same shape as min set to 1 and wondering why there's no redundancy.

The features that matter for Domain 2

Verbatim, because each answers a different exam question:

Health checks and replacement — "Amazon EC2 Auto Scaling automatically monitors the health and availability of your instances using EC2 health checks and replaces terminated or impaired instances to maintain your desired capacity." Plus "you can define custom health checks that are specific to your application to verify that it's responding as expected. If an instance fails your custom health check, it's automatically replaced."

This is the self-healing answer. If a stem says "unhealthy instances must be replaced automatically", that's an Auto Scaling group — even at a fixed size. A group with min = max = desired = 3 does no scaling at all and is still worth having, because it replaces failures.

Balancing across AZs — "You can specify multiple Availability Zones for your Auto Scaling group, and Amazon EC2 Auto Scaling balances your instances evenly across the Availability Zones as the group scales. This provides high availability and resiliency by protecting your applications from failures in a single location."

Recall SAA0 lesson 4: "A subnet must reside within a single Availability Zone." So "multiple Availability Zones" means multiple subnets on the group. Three AZs, three subnets.

Mixed instance types and purchase options — "Within a single Auto Scaling group, you can launch multiple instance types and purchase options (Spot and On-Demand Instances), allowing you to optimize costs through Spot Instance usage." This is a Domain 4 answer living in a Domain 2 service: a baseline of On-Demand plus Spot for the elastic portion.

Spot replacement — "If your group includes Spot Instances, Amazon EC2 Auto Scaling can automatically request replacement Spot capacity if your Spot Instances are interrupted. Through Capacity Rebalancing, Amazon EC2 Auto Scaling can also monitor and proactively replace your Spot Instances that are at an elevated risk of interruption."

Load balancer integration — "Whenever instances are launched or terminated, Amazon EC2 Auto Scaling automatically registers and deregisters the instances from the load balancer." You do not wire this up yourself, and an option that says you must is wrong.

Instance refresh — "provides a mechanism to update instances in a rolling fashion when you update your AMI or launch template. You can also use a phased approach, known as a canary deployment, to test a new AMI or launch template on a small set of instances before rolling it out." This is the immutable infrastructure answer that task statement 2.2 asks about: you don't patch instances, you replace them.

Lifecycle hooks — "useful for defining custom actions that are invoked as new instances launch or before instances are terminated." And for stateful workloads: "Lifecycle hooks also offer a mechanism for persisting state on shut down. … you can also use scale-in protection or custom termination policies to prevent instances with long-running processes from terminating early."

⚠️ Scale-in protection is the answer to "our batch jobs get killed mid-run when the fleet scales in." Not "turn off scaling".

Cost: "There are no additional fees with Amazon EC2 Auto Scaling … You only pay for the AWS resources … that you use."

Scaling policy types

⚠️ I did not retrieve the scaling-policy comparison (target tracking, step scaling, simple scaling, scheduled, predictive) from the page I fetched — it lists features, not policy types. Read docs.aws.amazon.com/autoscaling/ec2/userguide/scale-your-group.html before relying on the details.

What the exam guide itself gives you, under 2.1 skills, is "Determining scaling strategies for components used in an architecture design", and under 3.2 skills, "Identifying metrics and conditions to perform scaling actions". So the examinable shape is: what metric, and what threshold? A known daily pattern points at scheduled scaling; an unknown one points at tracking a metric.

Beyond EC2: the thing that scales best is the thing you don't run

The Well-Architected general design principle (General design principles, verified 2026-09-21) is the reasoning the exam rewards:

"Stop guessing your capacity needs: If you make a poor capacity decision when deploying a workload, you might end up sitting on expensive idle resources or dealing with the performance implications of limited capacity. With cloud computing, these problems can go away. You can use as much or as little capacity as you need, and scale in and out automatically."

So the scaling ladder, in order of decreasing operational overhead:

  1. Manual resizing — you guessed. Almost never the exam answer.
  2. EC2 Auto Scaling — you set the policy; AWS moves desired capacity.
  3. Managed service auto scaling — ECS Service Auto Scaling, DynamoDB auto scaling, Aurora replicas.
  4. Serverless — Lambda, Fargate, DynamoDB on-demand. Nothing to size.

⚠️ Auto Scaling is a control-plane activity, and the DR whitepaper says so explicitly in the context of failover: "Because Auto Scaling is a control plane activity, taking a dependency on it will lower the resiliency of your overall recovery strategy. It is a trade-off." (Disaster recovery options in the cloud, verified 2026-09-22)

That's lesson 6's territory, but note it now: the thing that makes your steady state cheap is the thing you shouldn't lean on during a disaster. AWS names the alternative — "provision sufficient capacity such that the recovery Region can handle the full production load as deployed. This statically stable configuration is called hot standby."

Static stability is the idea: a system that doesn't need to do anything — call an API, scale out, provision — in order to survive the failure.

Check yourself

  1. A legacy application keeps session state in instance memory and cannot be modified. What do you do, and why is it the second-best answer?
  2. An Auto Scaling group has min 2, max 4, desired 2. CPU hits 95% and stays there. Nothing scales past
    1. Is Auto Scaling broken?
  3. What does an Auto Scaling group with min = max = desired = 3 give you?
  4. Batch jobs are being killed mid-run when the group scales in. What's the fix?
  5. Why does AWS warn against depending on Auto Scaling during a failover?
Answers
  1. Sticky sessions (session affinity on the load balancer). It's second-best because losing an instance loses those users' sessions and the fleet unbalances — but the exam guide explicitly covers "applications not built for the cloud … when application changes are not possible", so it's the right answer when modification is off the table. Otherwise externalise session state to ElastiCache or DynamoDB.
  2. No. It's doing exactly what you configured: "ensures that your group never goes above this size." Raise max.
  3. Self-healing, not scaling. It "replaces terminated or impaired instances to maintain your desired capacity", and with multiple subnets it "balances your instances evenly across the Availability Zones." That's worth having even with no scaling policy at all.
  4. Scale-in protection, or a custom termination policy — AWS names both "to prevent instances with long-running processes from terminating early." Lifecycle hooks can also persist state on shutdown.
  5. Because "Auto Scaling is a control plane activity, taking a dependency on it will lower the resiliency of your overall recovery strategy." Control planes have lower availability design goals than data planes. The statically stable alternative is hot standby.

Teaching this section

← PreviousDecoupling — queues, topics, events and workflowsNext →Load balancing and multi-tier architecture