AWS Training
Modules Listen All tracks

← Design Resilient Architectures

Starts this lesson and continues through 17 more to the end of certification prep.

Load balancing and multi-tier architecture

Why this lesson matters

Task statement 2.1 lists "Load balancing concepts (for example, Application Load Balancer [ALB])" and "Multi-tier architectures". Task 2.2 lists load balancing again. And Domain 4's task 4.2 asks for the skill of "Determining an appropriate load balancing strategy (for example, Application Load Balancer [Layer 7] compared with Network Load Balancer [Layer 4] compared with Gateway Load Balancer)".

That last quote is the exam guide handing you the answer key: the load balancer question is an OSI layer question. Work out which layer the requirement lives at, and the type falls out.

What Elastic Load Balancing does for you

From What is Elastic Load Balancing?, verified 2026-09-22:

"Elastic Load Balancing automatically distributes your incoming traffic across multiple targets, such as EC2 instances, containers, and IP addresses, in one or more Availability Zones. It monitors the health of its registered targets, and routes traffic only to the healthy targets. Elastic Load Balancing scales your load balancer capacity automatically in response to changes in incoming traffic."

Three separate guarantees in that paragraph, and the exam tests each:

  1. Distribution across targets in one or more Availability Zones.
  2. Health checking — "routes traffic only to the healthy targets."
  3. Elasticity — the load balancer itself scales. You don't size it.

Plus: "You can also offload the work of encryption and decryption to your load balancer so that your compute resources can focus on their main work." That's TLS termination, and it's why ACM turns up here — from SAA1 lesson 5, certificates are Regional and CloudFront needs us-east-1.

The three current types

Application Load Balancer Network Load Balancer Gateway Load Balancer
OSI layer 7 — application 4 — transport 3 — network (per exam guide)
Verbatim "functions at the application layer, the seventh layer of the Open Systems Interconnection (OSI) model" "functions at the fourth layer of the Open Systems Interconnection (OSI) model" see ⚠️ below
Protocols HTTP / HTTPS "TCP, UDP, TCP_UDP, TLS, QUIC, and TCP_QUIC" GENEVE
Scale claim — "It can handle millions of requests per second" —
Static IP no yes — "one Elastic IP address per subnet" —
Routes on request content flow hash —
For web apps, microservices, containers extreme throughput, non-HTTP, static IP virtual appliances / inline inspection

⚠️ I did not fetch the Gateway Load Balancer user guide. Its row above comes from the exam guide's own Layer 3 framing plus SAA1 lesson 3's PrivateLink page, which describes a Gateway Load Balancer endpoint as sending "traffic to a fleet of virtual appliances using private IP addresses" routed via route tables. Read docs.aws.amazon.com/elasticloadbalancing/latest/gateway/ before relying on specifics.

⚠️ Classic Load Balancer is previous-generation. "We recommend that you migrate to a current generation load balancer." If it appears as an option, it is almost always the wrong answer.

ALB — routing on content

The mechanism, verbatim:

"After the load balancer receives a request, it evaluates the listener rules in priority order to determine which rule to apply, and then selects a target from the target group for the rule action."

Components, and the exam uses these words:

Routing algorithm: "The default routing algorithm is round robin; alternatively, you can specify the least outstanding requests routing algorithm."

What only an ALB can do

Verbatim from the benefits list — this is the list that makes an ALB the answer:

⚠️ "Lambda as an ALB target" surprises people. If a stem wants HTTP in front of a Lambda and rules out API Gateway on cost or simplicity grounds, an ALB is legitimate.

⚠️ ALB is the only one of the three that AWS WAF attaches to (with CloudFront, API Gateway REST, AppSync, Cognito user pool, App Runner, Bedrock AgentCore Gateway, Verified Access, Amplify — SAA1 lesson 4). NLB is not on that list. Layer 7 inspection requires a layer 7 device.

NLB — flow-based, and the static IP answer

"A Network Load Balancer functions at the fourth layer… It can handle millions of requests per second. After the load balancer receives a request from a client, it selects a target from a target group in the default action."

Target selection is a flow hash, and the detail matters:

"For TCP traffic, the load balancer selects a target using a flow hash algorithm based on the protocol, source IP address, source port, destination IP address, destination port, and TCP sequence number. … Each individual TCP connection is routed to a single target for the life of the connection."

"For UDP traffic… A UDP flow has the same source and destination, so it is consistently routed to a single target throughout its lifetime."

So an NLB does not read your request. It cannot route on a URL, because at layer 4 there is no URL.

The static IP property — the highest-yield NLB fact:

"Elastic Load Balancing creates a network interface for each Availability Zone you enable. Each load balancer node in the Availability Zone uses this network interface to get a static IP address. When you create an Internet-facing load balancer, you can optionally associate one Elastic IP address per subnet."

⚠️ "The firewall team needs a fixed IP to allowlist" → NLB. An ALB's addresses change; that's why you point at it with a Route 53 alias record (SAA0 lesson 4). If a stem mentions allowlisting, hardcoded IPs, or a client that can't do DNS, the answer is NLB with Elastic IPs.

Other NLB-specific benefits, verbatim: "Support for registering targets by IP address, including targets outside the VPC for the load balancer" — which covers on-premises targets over Direct Connect or VPN — and QUIC support "with advanced congestion control, fewer round trip connection establishment, built in TLS, and connection migration across networks."

Cross-zone load balancing — the trap

Both introductions describe the same node model:

"When you enable an Availability Zone for the load balancer, Elastic Load Balancing creates a load balancer node in the Availability Zone. By default, each load balancer node distributes traffic across the registered targets in its Availability Zone only. If you enable cross-zone load balancing, each load balancer node distributes traffic across the registered targets in all enabled Availability Zones."

⚠️ This is the classic "uneven distribution" question. Two AZs — one with 2 targets, one with 8. Traffic arrives roughly evenly at the two nodes, so each of the 2 targets gets ~25% and each of the 8 gets ~6%. The 2 melt.

Two fixes, and both are legitimate: enable cross-zone load balancing, or keep target counts balanced across AZs. AWS recommends the second independently:

"To increase the fault tolerance of your applications, you can enable multiple Availability Zones for your load balancer and ensure that each target group has at least one target in each enabled Availability Zone."

⚠️ I did not retrieve the per-type default for cross-zone load balancing, or its data-transfer charging. Both are examinable and commonly stated to differ between ALB and NLB — read docs.aws.amazon.com/elasticloadbalancing/latest/userguide/how-elastic-load-balancing-works.html and the pricing page.

The DNS failure mode, quoted because it is a real incident shape:

"if one or more target groups does not have a healthy target in an Availability Zone, we remove the IP address for the corresponding subnet from DNS, but the load balancer nodes in the other Availability Zones are still available to route traffic. If a client doesn't honor the time-to-live (TTL) and sends requests to the IP address after it is removed from DNS, the requests fail."

Clients that cache DNS forever keep hitting a withdrawn address. That's why NLB's static IPs are a double-edged tool, and why "the load balancer is healthy but some clients get errors" points at DNS caching.

Multi-tier: putting it together

Task 2.1 asks for "Multi-tier architectures" and 2.2 for "Implementing designs to mitigate single points of failure". The canonical answer, combining this lesson with SAA1 lesson 3 and SAA2 lesson 2:

  Route 53 alias ──▶ ALB (public subnets, ≥2 AZs)
                      │  SG: 443 from 0.0.0.0/0
                      ▼
                    Web/app tier — Auto Scaling group across ≥2 private subnets
                      │  SG: 443 from the ALB's security group
                      ▼
                    Data tier — RDS Multi-AZ in ≥2 private subnets
                         SG: 3306 from the app tier's security group

Every tier gets: its own subnets, in at least two AZs; its own security group, referencing the tier in front by group ID rather than CIDR; and no public route below the top tier.

⚠️ The load balancer needs a subnet in every AZ it serves. A subnet is zonal (SAA0 lesson 4), so "multi-AZ" always means "more subnets".

Global Accelerator is the layer above. From the ALB page: "Improves the availability and performance of your application. Use an accelerator to distribute traffic across multiple load balancers in one or more AWS Regions." The DR whitepaper adds that it "makes use of the extensive AWS edge network to put traffic on the AWS network backbone as soon as possible" and "avoids caching issues that can occur with DNS systems (like Route 53)."

⚠️ Global Accelerator vs Route 53 is a real exam pair. Route 53 is DNS — subject to client caching, and its failover speed depends on TTLs. Global Accelerator gives you static anycast IPs and moves traffic at the network layer, so it sidesteps DNS caching. If the stem complains that failover is slow because clients cache DNS, that's Global Accelerator.

Check yourself

  1. The requirement is to route /api/* and /static/* to different fleets. Which load balancer?
  2. A partner's firewall team needs two fixed IPs to allowlist. Which, and what do you attach?
  3. Two AZs: one has 2 instances, one has 8. The 2 are overloaded. What's happening and what are the two fixes?
  4. You need AWS WAF in front of a TCP-based game server behind an NLB. What do you tell the team?
  5. Failover takes 20 minutes because clients cache DNS. What changes?
Answers
  1. ALB. "Path conditions… forward requests based on the URL in the request." An NLB works at layer 4 where there is no URL to read.
  2. NLB, with one Elastic IP address per subnet — "you can optionally associate one Elastic IP address per subnet." ALB addresses change, which is why you reach it via a Route 53 alias instead.
  3. Cross-zone load balancing is off, so "each load balancer node distributes traffic across the registered targets in its Availability Zone only" — the 2 instances split the same share as the 8. Fix by enabling cross-zone load balancing, or by balancing target counts across AZs.
  4. It can't be done at the NLB. WAF attaches to CloudFront, API Gateway REST APIs, ALB, AppSync, Cognito user pools, App Runner, Bedrock AgentCore Gateway, Verified Access and Amplify — not NLB. And WAF inspects HTTP/HTTPS, so a TCP game protocol isn't WAF's problem anyway; that's security groups, NACLs and AWS Network Firewall.
  5. AWS Global Accelerator. Static anycast IPs, traffic moved at the network layer, and it "avoids caching issues that can occur with DNS systems (like Route 53)."

Teaching this section

← PreviousScaling — horizontal, vertical, and why state is the constraintNext →Containers and serverless — ECS, EKS, Fargate, Lambda