AWS Training
Modules Listen All tracks

← Design High-Performing Architectures

Starts this lesson and continues through 11 more to the end of certification prep.

Database performance — engines, replicas, proxies and caches

Why this lesson matters

Task statement 3.3 — "Determine high-performing database solutions" — has eight knowledge items and five skills, the most of any task in the domain. The item that organises all the others is "Data access patterns (for example, read-intensive compared with write-intensive)".

Every database performance fix helps one side of that split and does nothing for the other.

   READ-heavy                              WRITE-heavy
   ──────────                              ───────────
   read replicas / Aurora Replicas         bigger writer instance, Provisioned IOPS
   ElastiCache (lazy loading)              DynamoDB with enough WCUs / on-demand
   DAX (DynamoDB, eventually consistent)   shard / partition the keyspace
   CloudFront for cacheable responses      queue writes (SAA2 lesson 1)

A read replica does nothing for a write bottleneck. DAX is, in AWS's words, not ideal for "Applications that are write-intensive." Half the wrong answers in this area are right fixes for the other column.

Choosing the database type

The skill is "Determining an appropriate database type (for example, Amazon Aurora, Amazon DynamoDB)", and the knowledge item is "serverless, relational compared with non-relational, in-memory".

Need Type Service
joins, SQL, transactions across tables, existing MySQL/PostgreSQL app relational RDS or Aurora
key-value/document, consistent latency at any scale, no JOINs needed non-relational DynamoDB
sub-millisecond reads, cache or session store in-memory ElastiCache
unpredictable relational load, no instance sizing serverless relational Aurora Serverless

From What is Amazon DynamoDB?: "a serverless, fully managed, distributed NoSQL database with single-digit millisecond performance at any scale" — and ⚠️ "Unlike relational databases, DynamoDB doesn't support a JOIN operator. We recommend that you denormalize your data model." A requirement for ad-hoc joins across entities rules DynamoDB out.

⚠️ The in-scope list also names Amazon Redshift, DocumentDB, Keyspaces and Neptune. I did not fetch their documentation for this lesson, so it makes no claim about them. Read each service's "What is" page before relying on specifics.

Relational: RDS and Aurora

Aurora

From What is Amazon Aurora?, verified 2026-09-25:

"a fully managed relational database engine that's compatible with MySQL and PostgreSQL. … Aurora delivers up to 6x the throughput of stock PostgreSQL and up to 6x the throughput of stock MySQL on similar hardware."

⚠️ Older prep material quotes different multiples. The page says 6x for both, as of 2026-09-25.

*"The underlying storage grows automatically as needed. An Aurora cluster volume can grow to a maximum size of 256 tebibytes (TiB)."*

Aurora Replicas

From Replication with Amazon Aurora:

⚠️ Aurora Replicas do both jobs; RDS does not. The RDS read-replica page says it directly: "Unlike a read replica, a standby replica can't serve read traffic." In Aurora, the same replicas scale reads and are the failover targets. That's a real design difference, and the exam tests it.

Cross-Region is engine-dependent, verbatim: Aurora MySQL can use an Aurora global database or binlog replicas ("up to five read replicas created this way, each in a different Region"). Aurora PostgreSQL "doesn't support cross-Region Aurora Replicas" — use a global database, "up to 10 read-only secondary DB clusters in different Regions."

Aurora Serverless

From Using Aurora serverless: "an on-demand, autoscaling configuration", "especially valuable for multitenant databases, distributed databases, development and test systems, and other environments with highly variable and unpredictable workloads." It scales in half-ACU steps — "It can add 0.5, 1, 1.5, 2, or additional half-ACUs" — and "Scaling can happen while SQL statements are running and transactions are open." Reader instances can also be serverless and "scale independently of the writer".

⚠️ The page I fetched does not state the ACU minimum and maximum or the memory per ACU. It links to "Scaling to Zero ACUs with automatic pause and resume", which implies a zero minimum is possible, but read docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/aurora-serverless-v2.setting-capacity.html before quoting the range.

RDS read replicas

From Working with DB instance read replicas:

"you can elastically scale out beyond the capacity constraints of a single DB instance for read-heavy database workloads."

⚠️ The pages I fetched do not state the maximum number of read replicas per source instance for each engine. Read the per-engine replica pages linked from USER_ReadRepl.html before relying on a number.

MySQL vs PostgreSQL — what the replica page actually distinguishes

The skill "Determining an appropriate database engine (for example, MySQL compared with PostgreSQL)" is thinly documented as a comparison. What I can state from Differences between read replicas for DB engines:

RDS for MySQL / MariaDB RDS for PostgreSQL
Replication method Logical Physical
Writable replicas "You can enable the MariaDB or MySQL read replica to be writable." "PostgreSQL doesn't allow for a read replica to be made writable."
Parallel replication "allow for parallel replication threads" "a single process handling replication"
Replica of a replica supported "certain versions"

Plus the Aurora cross-Region difference above: Aurora MySQL has binlog cross-Region replicas; Aurora PostgreSQL uses global databases only.

⚠️ Beyond replication, no page I fetched compares the two engines' features. In practice the exam's engine-choice questions usually turn on compatibility with an existing application — keep the same engine to avoid a migration — which is the next section.

Heterogeneous vs homogeneous migrations

The knowledge item is "Database engines with appropriate use cases (for example, heterogeneous migrations, homogeneous migrations)". From What is AWS DMS?:

→ Engine changes ⇒ schema conversion first (DMS Schema Conversion or SCT), then DMS. Same engine ⇒ no schema conversion. Minimal downtime ⇒ DMS with ongoing replication, cut over when in sync.

Capacity planning — instance, storage, capacity units

The knowledge item is "Database capacity planning (for example, capacity units, instance types, Provisioned IOPS)".

RDS storage and Provisioned IOPS

From Amazon RDS DB instance storage:

Provisioned IOPS (io2 Block Express) General Purpose (gp3)
AWS's framing "designed to meet the needs of I/O-intensive workloads, particularly database workloads, that require low I/O latency and consistent I/O throughput" — "best suited for production" "cost-effective storage that is ideal for a broad range of workloads running on medium-sized DB instances" — "best suited for development and testing"
Max IOPS 256,000 64,000 (16,000 on SQL Server)
Latency "Sub-millisecond, provided consistently 99.9% of the time" "Single-digit millisecond, provided consistently 99% of the time"
Baseline — "3000 IOPS and 125 MiB/s"; at 400 GiB+ (most engines) "12,000 IOPS and 500 MiB/s" via volume striping

⚠️ The instance caps the storage. "if the maximum IOPS for a particular DB instance class is 40,000, and you attach four 64,000 IOPS EBS volumes, the maximum IOPS is 40,000 rather than 256,000." And AWS says to "determine the maximum bandwidth, throughput and IOPS for the instance class before setting a Provisioned IOPS". The metric to watch: DiskQueueDepth — "I/O requests in the queue waiting to be serviced". If TotalIOPS "regularly approach the Provisioned IOPS value", raise it.

DynamoDB capacity units

From DynamoDB provisioned capacity mode:

"For an item up to 4 KB, one read capacity unit (RCU) represents one strongly consistent read operation per second, or two eventually consistent read operations per second."

"A write capacity unit (WCU) represents one write per second for an item up to 1 KB."

Round item sizes up. AWS's worked examples: 80 strongly consistent reads/s of 3 KB items → 3 ÷ 4 = 0.75 → 1 RCU each → 80 RCU. 100 writes/s of 512-byte items → 100 WCU. Eventually consistent reads halve the RCU requirement — and "Read operations in DynamoDB are eventually consistent, by default."

On-demand vs provisioned. On-demand "offers pay-as-you-go pricing", "instantly scales up or down" and "scales down to zero". Provisioned charges "based on the hourly read and write capacity you have provisioned, not how much of that provisioned capacity you actually consumed", and can use auto scaling — AWS recommends "setting target utilization to 70%". The default account table-level quota is 40,000 read and write throughput units, adjustable.

→ Unpredictable/spiky/new → on-demand. Steady and forecastable → provisioned (+ auto scaling).

Connections and proxies — RDS Proxy

SAA2 lesson 5 covered RDS Proxy for resilience. The performance angle, from Amazon RDS Proxy:

"RDS Proxy establishes a database connection pool and reuses connections in this pool. This approach avoids the memory and CPU overhead of opening a new database connection each time."

It handles "unpredictable surges in database traffic" that would otherwise cause "oversubscribing connections or new connections being created at a fast rate". Enable it "for most applications with no code changes."

⚠️ Limits that eliminate options: "you can associate a proxy only with the writer DB instance, not a read replica"; the proxy "must be in the same virtual private cloud (VPC) as the database" and "can't be publicly accessible"; 20 proxies per account.

→ Lambda (or any fast-scaling fleet) + connection errors on RDS = RDS Proxy. It doesn't make queries faster; it stops the connection count from being the bottleneck.

Caching — ElastiCache

The knowledge item is "Caching strategies and services (for example, Amazon ElastiCache)". From What is Amazon ElastiCache?: "a distributed in-memory data store or cache", working with Valkey, Memcached and Redis OSS, in two formats — serverless ("eliminates the need to provision instances or configure nodes", "a highly available cache in under a minute") or node-based (you choose node type, count and AZ placement).

Memcached vs Valkey/Redis OSS

From Comparing node-based Valkey, Memcached, and Redis OSS clusters:

Memcached Valkey / Redis OSS
Data types simple complex (sets, sorted sets, lists, hashes…)
Multi-threaded yes no
Replication / high availability no yes
Automatic failover no optional (cluster mode disabled) / required (enabled)
Pub/Sub, sorted sets, geospatial no yes
Backup and restore serverless only yes

AWS's "choose Memcached if": "You need the simplest model possible", "large nodes with multiple cores or threads", "scale out and in", "cache objects".

→ Leaderboard (sorted sets), pub/sub, replication or failover ⇒ Valkey/Redis OSS. Simplest multi-threaded object cache ⇒ Memcached.

The two strategies

From Caching strategies:

Lazy loading Write-through
How read cache → on miss, read DB → write to cache on every DB write, also write cache
Advantage "Only requested data is cached"; "Node failures aren't fatal" "Data in the cache is never stale"
Disadvantage cache miss penalty — "three trips"; stale data missing data on a new node; cache churn — "Most data is never read"

And the combination: "By adding a time to live (TTL) value to each write, you can have the advantages of each strategy."

→ "Stale reads are unacceptable" ⇒ write-through (+ TTL). "Cache only what's actually read" ⇒ lazy loading (+ TTL).

Caching — DAX, only for DynamoDB

From In-memory acceleration with DynamoDB Accelerator (DAX): DAX "reduces the response times of eventually consistent read workloads by an order of magnitude from single-digit milliseconds to microseconds", and it's "API-compatible with DynamoDB" so it needs "only minimal functional changes". The Introduction page puts it as "up to 10 times performance improvement – from milliseconds to microseconds".

DAX is not ideal for, verbatim:

→ DynamoDB + microseconds + read-heavy + eventual consistency OK ⇒ DAX. Any of those three "not ideal" conditions ⇒ not DAX.

Global infrastructure

The knowledge item "AWS global infrastructure (for example, Availability Zones, AWS Regions)" is about placement. DynamoDB "automatically replicates your data across three Availability Zones" with a "99.99% availability SLA"; global tables give "multi-active replication" with a "99.999% availability SLA" and "local read and write performance" in each Region. Aurora global databases give read-only secondaries in other Regions. For performance, the point is locality: put the read copy in the Region where the users are. (The DR use of both is SAA2 lesson 6.)

Choosing, under exam conditions

The stem says Answer
"reporting queries slow down the production MySQL database" read replica, point reporting at it
"Aurora cluster, read traffic growing, need automatic scaling of readers" Aurora Replicas + reader endpoint + AWS Auto Scaling
"RDS I/O latency high, DiskQueueDepth climbing, production OLTP" Provisioned IOPS (io2 Block Express), check instance class limit
"Lambda functions exhaust RDS connections" RDS Proxy
"sub-ms reads of hot product data in front of RDS" ElastiCache (lazy loading + TTL)
"reads must never be stale" write-through caching
"leaderboard with ranking" ElastiCache for Valkey/Redis OSS (sorted sets)
"DynamoDB reads need microseconds, read-heavy" DAX
"DynamoDB, strongly consistent reads, write-heavy" not DAX — capacity (WCU / on-demand)
"traffic unpredictable, new app, DynamoDB" on-demand capacity
"Oracle to Aurora PostgreSQL, minimal downtime" DMS Schema Conversion (or SCT) + DMS with ongoing replication
"relational workload, spiky and unpredictable, avoid instance sizing" Aurora Serverless

Check yourself

  1. A table reads 3 KB items at 100 strongly consistent reads per second. How many RCUs? And eventually consistent?
  2. Why is DAX wrong for a banking ledger that requires strongly consistent reads?
  3. Name two things Aurora Replicas do that an RDS Multi-AZ standby doesn't.
  4. What do you migrate first when moving from SQL Server to Aurora MySQL?
  5. Can you attach RDS Proxy to a read replica?
Answers
  1. 3 KB rounds up to one 4 KB unit → 100 RCU strongly consistent; eventually consistent is two reads per RCU → 50 RCU.
  2. DAX serves "eventually consistent data" and is not ideal for "Applications that require strongly consistent reads".
  3. They serve reads (through the reader endpoint) and are failover targets; there can be up to 15 of them. For RDS Multi-AZ: "Unlike a read replica, a standby replica can't serve read traffic."
  4. The schema — heterogeneous migration, so DMS Schema Conversion (or AWS SCT) first, then DMS for the data.
  5. No. "you can associate a proxy only with the writer DB instance, not a read replica."

Teaching this section

← PreviousCompute performance — instance families, scaling signals, Batch, Lambda memoryNext →Network performance — edge, hybrid links, placement