AWS Training
Modules Listen All tracks

← Design High-Performing Architectures

Starts this lesson and continues through 12 more to the end of certification prep.

Compute performance — instance families, scaling signals, Batch, Lambda memory

Why this lesson matters

Task statement 3.2 — "Design high-performing and elastic compute solutions" — has the longest knowledge list in the domain, and half of it repeats SAA2: "Queuing and messaging concepts", "Serverless technologies and patterns (for example, AWS Lambda, Fargate)", "The orchestration of containers (for example, Amazon ECS, Amazon EKS)". Those are taught in SAA2 lessons 1, 2 and 4, and this lesson doesn't repeat them.

What's new here is the sizing and signalling half. The four skills:

And three named services SAA2 didn't cover: "AWS Batch, Amazon EMR, AWS Fargate" as compute choices, plus "AWS Auto Scaling" as distinct from EC2 Auto Scaling.

The question this lesson answers is: once you've picked the compute model, how big, how many, and driven by what signal?

Reading an instance type name

From Amazon EC2 instance type naming conventions, verified 2026-09-25:

"The first position of the instance family indicates the series, for example c. The second position indicates the generation, for example 7. The third position indicates the options, for example gn. After the period (.) is the instance size."

   c   7   gn  .  xlarge
   │   │   │      └── size
   │   │   └── options: g = Graviton, n = network and EBS optimized
   │   └── generation
   └── series: C = compute optimized

The series letters worth knowing cold, verbatim:

Letter Meaning Exam signal
M General purpose "balanced", web/app servers
T Burstable performance dev/test, low average CPU with spikes
C Compute optimized "CPU-bound", batch, encoding
R Memory optimized in-memory caches, large DB buffers
X / U / Z Memory intensive / High memory / High memory SAP-scale memory
I / D Storage optimized / Dense storage high local I/O, instance store
P / G GPU accelerated / Graphics intensive ML training, rendering
Inf / Trn AWS Inferentia / AWS Trainium ML inference / training
Hpc High performance computing tightly coupled HPC

And the option letters that change an answer: g "AWS Graviton processors", a "AMD processors", d "Instance store volumes", n "Network and EBS optimized", e extra storage or memory, flex "Flex instance".

Burstable and flex — the CPU-credit trap

From Amazon EC2 instance types:

Burstable performance (T) instances "provide a baseline level of CPU performance with the ability to burst above the baseline", governed by CPU credits. Fixed performance instances "can deliver and sustain full CPU performance at any time … If you need consistently high CPU performance for applications such as video encoding, high volume websites, or HPC applications, we recommend that you use fixed performance instances."

Flex instances deliver "a baseline CPU performance of 40 percent" and can go "up to 100 percent CPU performance for 95 percent of the time over a 24-hour window."

⚠️ The pattern is the gp2 trap from lesson 1, on CPU. "A t3 instance performs well, then slows under sustained load" is credit exhaustion. The fix for sustained high CPU is a fixed-performance family (M or C), not a bigger T.

Scaling signals — which metric, and why

SAA2 lesson 2 covered min/max/desired and self-healing. The skill here is "Identifying metrics and conditions to perform scaling actions".

From Dynamic scaling for Amazon EC2 Auto Scaling, verified 2026-09-25, the three dynamic policy types:

Policy What it does
Target tracking "based on a Amazon CloudWatch metric and a target value" — "similar to the way that your thermostat maintains the temperature"
Step scaling "step adjustments, that vary based on the size of the alarm breach"
Simple scaling "a single scaling adjustment, with a cooldown period between each scaling activity"

And the rule that decides which metric:

"We strongly recommend that you use target tracking scaling policies and choose a metric that changes inversely proportional to a change in the capacity of your Auto Scaling group. So if you double the size of your Auto Scaling group, the metric decreases by 50 percent."

AWS's examples: "average CPU utilization or average request count per target."

⚠️ Why queue length is the wrong metric. From Scaling policy based on Amazon SQS: "the number of messages in the queue might not change proportionally to the size of the Auto Scaling group that processes messages from the queue." The fix is backlog per instance:

   backlog per instance       = ApproximateNumberOfMessages ÷ InService instances
   acceptable backlog (target) = acceptable latency ÷ average processing time per message

AWS's worked example: 10 instances, 1,500 messages, 0.1 s per message, 10 s acceptable latency → target = 10 ÷ 0.1 = 100. Current backlog = 1,500 ÷ 10 = 150, so the group scales out "by five instances to maintain proportion to the target value". The page now recommends computing this with metric math instead of publishing a custom metric.

This is also the answer to the first skill — "Decoupling workloads so that components can scale independently". A queue between tiers lets the worker tier scale on its own backlog, not on the web tier's CPU.

Multiple policies: "Amazon EC2 Auto Scaling chooses the policy that provides the largest capacity for both scale out and scale in." The intention "is to prevent Amazon EC2 Auto Scaling from removing too many instances."

⚠️ I did not fetch the scheduled-scaling or predictive-scaling pages for EC2 Auto Scaling. The scaling plans page (below) confirms predictive scaling exists "to scale your Amazon EC2 capacity faster when there are regularly occurring spikes". For scheduled scaling semantics, read docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-scheduled-scaling.html.

AWS Auto Scaling vs EC2 Auto Scaling

The guide names both. From What is a scaling plan?, AWS Auto Scaling scaling plans "configure auto scaling for related or associated scalable resources" across:

→ One service's instances = EC2 Auto Scaling. Several resource types scaled together as an application = AWS Auto Scaling. AWS also notes: "If you use scaling plans only for predictive scaling, we strongly recommend that you set predictive scaling policies directly on your Auto Scaling resources instead."

AWS Batch — queued jobs, not a fleet you manage

From What is AWS Batch?:

"AWS Batch automatically provisions compute resources and optimizes the workload distribution based on the quantity and scale of the workloads. With AWS Batch, there's no need to install or manage batch computing software."

It runs "on top of AWS managed container orchestration services, Amazon ECS and Amazon EKS", and scales "Amazon EC2 instances, Fargate resources, and Amazon ECS Managed Instances" — including "Spot or On-Demand Instances".

The components, from Components of AWS Batch:

Component What it is
Job "A unit of work (such as a shell script, a Linux executable, or a Docker container image)" — can depend on other jobs
Job definition "a blueprint for the resources in your job" — IAM role, memory and CPU
Job queue where the job "resides until it's scheduled onto a compute environment"; queues have priority
Compute environment managed (you pick Fargate, EC2 or ECS Managed Instances, min/desired/max vCPUs, Spot price threshold) or unmanaged
Scheduling policy fair-share; default is "first-in, first-out (FIFO)"

AWS's own example of priority: "a high priority queue that you submit time-sensitive jobs to, and a low priority queue for jobs that can run anytime when compute resources are cheaper."

→ Batch is the answer when the stem describes jobs — thousands of them, queued, with dependencies, and a cost target that says Spot — not a long-running service. If the job needs more than Lambda's 15-minute ceiling (SAA2 lesson 4), Batch is usually the managed alternative.

Fargate — sizing is per task

From Architect for AWS Fargate for Amazon ECS: with Fargate you "specify the CPU and memory requirements" per task, and "Each Fargate task has its own isolation boundary and does not share the underlying kernel, CPU resources, memory resources, or elastic network interface with another task." Fargate Spot runs "interruption tolerant" tasks at a discount with "a two-minute warning".

⚠️ I did not retrieve the table of valid Fargate CPU/memory combinations; the Fargate page links to it. Read docs.aws.amazon.com/AmazonECS/latest/developerguide/fargate-tasks-services.html#fargate-tasks-size before quoting a maximum task size.

Lambda — memory is the CPU knob

From Configure Lambda function memory, verified 2026-09-25:

"Lambda allocates CPU power in proportion to the amount of memory configured. … You can configure memory between 128 MB and 10,240 MB in 1-MB increments. At 1,769 MB, a function has the equivalent of one vCPU."

And: "Memory is the principal lever for controlling the performance of a function." "If a function is CPU, network or memory-bound, then increasing the memory setting can dramatically improve its performance." The default, 128 MB, is "the lowest possible setting" and AWS recommends it only "for simple Lambda functions, such as those that transform and route events."

⚠️ There is no separate CPU setting. A stem saying "the Lambda function is CPU-bound; increase its CPU" has exactly one mechanism: raise memory. And because Lambda bills per request plus GB-seconds (SAA2 lesson 4), a CPU-bound function can cost the same or less at higher memory if it finishes proportionally sooner — the lab measures this rather than assuming it.

Tools AWS names for finding the right size: CloudWatch (alarm when memory approaches the maximum; watch duration for CPU- and I/O-bound functions), the open-source AWS Lambda Power Tuning tool ("uses AWS Step Functions to run multiple concurrent versions of a Lambda function at different memory allocations"), and AWS Compute Optimizer recommendations — which "supports only functions that use x86_64 architecture."

EMR as a compute choice

The guide lists EMR under 3.2 as a compute service, and again under 3.5 for data processing (lesson 6). The compute-sizing facts, from What is Amazon EMR? and Understanding how to create and work with Amazon EMR clusters:

Node type Role
Primary "manages the cluster"; every cluster has one
Core "run tasks and store data in the Hadoop Distributed File System (HDFS)"
Task "only runs tasks and does not store data in HDFS. Task nodes are optional."

⚠️ This is the Spot question. Task nodes hold no HDFS data, so losing one loses no data. That makes task nodes the natural place for Spot capacity; core nodes hold HDFS and are riskier to lose. Instance fleets can mix "target capacities for On-Demand and Spot Instances".

And the cluster lifecycle choice: a transient cluster runs its steps and terminates ("typically done for clusters that process a set amount of data and then terminate"), versus a long-running cluster you submit steps to. Note: "If a cluster terminates because of a failure, any data stored on the cluster is deleted" — which is why EMR output goes to S3.

Distributed computing and the edge

The knowledge item "Distributed computing concepts supported by AWS global infrastructure and edge services" is served by CloudFront and Global Accelerator, which lesson 4 covers.

⚠️ I did not fetch the Lambda@Edge or CloudFront Functions pages, so this lesson makes no claim about running compute at edge locations. Read docs.aws.amazon.com/AmazonCloudFront/latest/DeveloperGuide/edge-functions.html.

Choosing, under exam conditions

The stem says Answer
"CPU-bound, consistently high utilisation" C family, fixed performance
"large in-memory dataset, cache nodes" R family
"t3 slows down after sustained load" CPU credits → M or C (fixed performance)
"scale web tier on load, simplest correct policy" target tracking on CPU or request count per target
"scale workers on an SQS queue" target tracking on backlog per instance
"scale ASG, ECS tasks, DynamoDB and Aurora replicas together" AWS Auto Scaling (scaling plan)
"thousands of containerised jobs, dependencies, cheapest capacity" AWS Batch with Spot compute environment
"Lambda function is CPU-bound" increase memory (1,769 MB = one vCPU)
"Hadoop/Spark cluster, cut cost without risking HDFS data" Spot on task nodes
"job runs once a night, then nothing" transient EMR cluster, output to S3

Check yourself

  1. Decode r7gd.2xlarge.
  2. Why does AWS say queue length is the wrong target-tracking metric, and what replaces it?
  3. How do you give a Lambda function more CPU?
  4. Which EMR node type is safest to run on Spot, and why?
  5. What's the difference between EC2 Auto Scaling and AWS Auto Scaling in one sentence?
Answers
  1. R = memory optimized, 7 = generation, g = Graviton, d = instance store volumes, 2xlarge = size.
  2. The queue length "might not change proportionally to the size of the Auto Scaling group". Use backlog per instance, with target = acceptable latency ÷ average processing time.
  3. Raise its memory. "Lambda allocates CPU power in proportion to the amount of memory configured"; at 1,769 MB it has one vCPU.
  4. Task nodes — they "only runs tasks and does not store data in HDFS", so an interruption loses no data.
  5. EC2 Auto Scaling scales one Auto Scaling group's instances; AWS Auto Scaling scaling plans scale several resource types together — ASGs, ECS tasks, DynamoDB capacity, Aurora replicas, Spot Fleets.

Teaching this section

← PreviousStorage performance — object, file, block, and the disk that disappearsNext →Database performance — engines, replicas, proxies and caches