Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Starts this lesson and continues through 12 more to the end of certification prep.
Task statement 3.2 — "Design high-performing and elastic compute solutions" — has the longest
knowledge list in the domain, and half of it repeats SAA2: "Queuing and messaging concepts",
"Serverless technologies and patterns (for example, AWS Lambda, Fargate)", "The orchestration of
containers (for example, Amazon ECS, Amazon EKS)". Those are taught in SAA2 lessons 1, 2 and 4, and
this lesson doesn't repeat them.
What's new here is the sizing and signalling half. The four skills:
And three named services SAA2 didn't cover: "AWS Batch, Amazon EMR, AWS Fargate" as compute
choices, plus "AWS Auto Scaling" as distinct from EC2 Auto Scaling.
The question this lesson answers is: once you've picked the compute model, how big, how many, and driven by what signal?
From Amazon EC2 instance type naming conventions, verified 2026-09-25:
"The first position of the instance family indicates the series, for example
c. The second position indicates the generation, for example7. The third position indicates the options, for examplegn. After the period (.) is the instance size."
c 7 gn . xlarge
│ │ │ └── size
│ │ └── options: g = Graviton, n = network and EBS optimized
│ └── generation
└── series: C = compute optimized
The series letters worth knowing cold, verbatim:
| Letter | Meaning | Exam signal |
|---|---|---|
| M | General purpose | "balanced", web/app servers |
| T | Burstable performance | dev/test, low average CPU with spikes |
| C | Compute optimized | "CPU-bound", batch, encoding |
| R | Memory optimized | in-memory caches, large DB buffers |
| X / U / Z | Memory intensive / High memory / High memory | SAP-scale memory |
| I / D | Storage optimized / Dense storage | high local I/O, instance store |
| P / G | GPU accelerated / Graphics intensive | ML training, rendering |
| Inf / Trn | AWS Inferentia / AWS Trainium | ML inference / training |
| Hpc | High performance computing | tightly coupled HPC |
And the option letters that change an answer: g "AWS Graviton processors", a "AMD
processors", d "Instance store volumes", n "Network and EBS optimized", e extra
storage or memory, flex "Flex instance".
From Amazon EC2 instance types:
Burstable performance (
T) instances "provide a baseline level of CPU performance with the ability to burst above the baseline", governed by CPU credits. Fixed performance instances "can deliver and sustain full CPU performance at any time … If you need consistently high CPU performance for applications such as video encoding, high volume websites, or HPC applications, we recommend that you use fixed performance instances."
Flex instances deliver "a baseline CPU performance of 40 percent" and can go "up to 100 percent CPU performance for 95 percent of the time over a 24-hour window."
⚠️ The pattern is the gp2 trap from lesson 1, on CPU. "A t3 instance performs well, then slows
under sustained load" is credit exhaustion. The fix for sustained high CPU is a fixed-performance
family (M or C), not a bigger T.
SAA2 lesson 2 covered min/max/desired and self-healing. The skill here is "Identifying metrics and
conditions to perform scaling actions".
From Dynamic scaling for Amazon EC2 Auto Scaling, verified 2026-09-25, the three dynamic policy types:
| Policy | What it does |
|---|---|
| Target tracking | "based on a Amazon CloudWatch metric and a target value" — "similar to the way that your thermostat maintains the temperature" |
| Step scaling | "step adjustments, that vary based on the size of the alarm breach" |
| Simple scaling | "a single scaling adjustment, with a cooldown period between each scaling activity" |
And the rule that decides which metric:
"We strongly recommend that you use target tracking scaling policies and choose a metric that changes inversely proportional to a change in the capacity of your Auto Scaling group. So if you double the size of your Auto Scaling group, the metric decreases by 50 percent."
AWS's examples: "average CPU utilization or average request count per target."
⚠️ Why queue length is the wrong metric. From Scaling policy based on Amazon SQS: "the number of messages in the queue might not change proportionally to the size of the Auto Scaling group that processes messages from the queue." The fix is backlog per instance:
backlog per instance = ApproximateNumberOfMessages ÷ InService instances
acceptable backlog (target) = acceptable latency ÷ average processing time per message
AWS's worked example: 10 instances, 1,500 messages, 0.1 s per message, 10 s acceptable latency → target = 10 ÷ 0.1 = 100. Current backlog = 1,500 ÷ 10 = 150, so the group scales out "by five instances to maintain proportion to the target value". The page now recommends computing this with metric math instead of publishing a custom metric.
This is also the answer to the first skill — "Decoupling workloads so that components can scale independently". A queue between tiers lets the worker tier scale on its own backlog, not on the web tier's CPU.
Multiple policies: "Amazon EC2 Auto Scaling chooses the policy that provides the largest capacity for both scale out and scale in." The intention "is to prevent Amazon EC2 Auto Scaling from removing too many instances."
⚠️ I did not fetch the scheduled-scaling or predictive-scaling pages for EC2 Auto Scaling. The scaling
plans page (below) confirms predictive scaling exists "to scale your Amazon EC2 capacity faster when
there are regularly occurring spikes". For scheduled scaling semantics, read
docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-scheduled-scaling.html.
The guide names both. From What is a scaling plan?, AWS Auto Scaling scaling plans "configure auto scaling for related or associated scalable resources" across:
→ One service's instances = EC2 Auto Scaling. Several resource types scaled together as an application = AWS Auto Scaling. AWS also notes: "If you use scaling plans only for predictive scaling, we strongly recommend that you set predictive scaling policies directly on your Auto Scaling resources instead."
From What is AWS Batch?:
"AWS Batch automatically provisions compute resources and optimizes the workload distribution based on the quantity and scale of the workloads. With AWS Batch, there's no need to install or manage batch computing software."
It runs "on top of AWS managed container orchestration services, Amazon ECS and Amazon EKS", and scales "Amazon EC2 instances, Fargate resources, and Amazon ECS Managed Instances" — including "Spot or On-Demand Instances".
The components, from Components of AWS Batch:
| Component | What it is |
|---|---|
| Job | "A unit of work (such as a shell script, a Linux executable, or a Docker container image)" — can depend on other jobs |
| Job definition | "a blueprint for the resources in your job" — IAM role, memory and CPU |
| Job queue | where the job "resides until it's scheduled onto a compute environment"; queues have priority |
| Compute environment | managed (you pick Fargate, EC2 or ECS Managed Instances, min/desired/max vCPUs, Spot price threshold) or unmanaged |
| Scheduling policy | fair-share; default is "first-in, first-out (FIFO)" |
AWS's own example of priority: "a high priority queue that you submit time-sensitive jobs to, and a low priority queue for jobs that can run anytime when compute resources are cheaper."
→ Batch is the answer when the stem describes jobs — thousands of them, queued, with dependencies,
and a cost target that says Spot — not a long-running service. If the job needs more than Lambda's
15-minute ceiling (SAA2 lesson 4), Batch is usually the managed alternative.
From Architect for AWS Fargate for Amazon ECS: with Fargate you "specify the CPU and memory requirements" per task, and "Each Fargate task has its own isolation boundary and does not share the underlying kernel, CPU resources, memory resources, or elastic network interface with another task." Fargate Spot runs "interruption tolerant" tasks at a discount with "a two-minute warning".
⚠️ I did not retrieve the table of valid Fargate CPU/memory combinations; the Fargate page links
to it. Read docs.aws.amazon.com/AmazonECS/latest/developerguide/fargate-tasks-services.html#fargate-tasks-size
before quoting a maximum task size.
From Configure Lambda function memory, verified 2026-09-25:
"Lambda allocates CPU power in proportion to the amount of memory configured. … You can configure memory between 128 MB and 10,240 MB in 1-MB increments. At 1,769 MB, a function has the equivalent of one vCPU."
And: "Memory is the principal lever for controlling the performance of a function." "If a function is CPU, network or memory-bound, then increasing the memory setting can dramatically improve its performance." The default, 128 MB, is "the lowest possible setting" and AWS recommends it only "for simple Lambda functions, such as those that transform and route events."
⚠️ There is no separate CPU setting. A stem saying "the Lambda function is CPU-bound; increase its
CPU" has exactly one mechanism: raise memory. And because Lambda bills per request plus GB-seconds
(SAA2 lesson 4), a CPU-bound function can cost the same or less at higher memory if it finishes
proportionally sooner — the lab measures this rather than assuming it.
Tools AWS names for finding the right size: CloudWatch (alarm when memory approaches the maximum; watch duration for CPU- and I/O-bound functions), the open-source AWS Lambda Power Tuning tool ("uses AWS Step Functions to run multiple concurrent versions of a Lambda function at different memory allocations"), and AWS Compute Optimizer recommendations — which "supports only functions that use x86_64 architecture."
The guide lists EMR under 3.2 as a compute service, and again under 3.5 for data processing (lesson 6). The compute-sizing facts, from What is Amazon EMR? and Understanding how to create and work with Amazon EMR clusters:
| Node type | Role |
|---|---|
| Primary | "manages the cluster"; every cluster has one |
| Core | "run tasks and store data in the Hadoop Distributed File System (HDFS)" |
| Task | "only runs tasks and does not store data in HDFS. Task nodes are optional." |
⚠️ This is the Spot question. Task nodes hold no HDFS data, so losing one loses no data. That makes task nodes the natural place for Spot capacity; core nodes hold HDFS and are riskier to lose. Instance fleets can mix "target capacities for On-Demand and Spot Instances".
And the cluster lifecycle choice: a transient cluster runs its steps and terminates ("typically done for clusters that process a set amount of data and then terminate"), versus a long-running cluster you submit steps to. Note: "If a cluster terminates because of a failure, any data stored on the cluster is deleted" — which is why EMR output goes to S3.
The knowledge item "Distributed computing concepts supported by AWS global infrastructure and edge services" is served by CloudFront and Global Accelerator, which lesson 4 covers.
⚠️ I did not fetch the Lambda@Edge or CloudFront Functions pages, so this lesson makes no claim about
running compute at edge locations. Read
docs.aws.amazon.com/AmazonCloudFront/latest/DeveloperGuide/edge-functions.html.
| The stem says | Answer |
|---|---|
| "CPU-bound, consistently high utilisation" | C family, fixed performance |
| "large in-memory dataset, cache nodes" | R family |
"t3 slows down after sustained load" |
CPU credits → M or C (fixed performance) |
| "scale web tier on load, simplest correct policy" | target tracking on CPU or request count per target |
| "scale workers on an SQS queue" | target tracking on backlog per instance |
| "scale ASG, ECS tasks, DynamoDB and Aurora replicas together" | AWS Auto Scaling (scaling plan) |
| "thousands of containerised jobs, dependencies, cheapest capacity" | AWS Batch with Spot compute environment |
| "Lambda function is CPU-bound" | increase memory (1,769 MB = one vCPU) |
| "Hadoop/Spark cluster, cut cost without risking HDFS data" | Spot on task nodes |
| "job runs once a night, then nothing" | transient EMR cluster, output to S3 |
r7gd.2xlarge.