AWS Training
Modules Listen All tracks

← Design High-Performing Architectures

Starts this lesson and continues through 9 more to the end of certification prep.

Ingestion and streaming — Kinesis, Firehose, and moving data in

Why this lesson matters

Task statement 3.5 — "Determine high-performing data ingestion and transformation solutions" — is the widest task in the domain: seven knowledge items and seven skills. This lesson takes the ingestion half; lesson 6 takes transformation and analytics.

The ingestion items, verbatim:

And the skills "Designing data streaming architectures", "Designing data transfer solutions", and "Selecting appropriate configurations for ingestion".

"Frequency" is the organising word. How often does the data arrive?

   CONTINUOUSLY, record by record, seconds matter    →  STREAM    →  Kinesis Data Streams / Firehose
   ON A SCHEDULE, files or datasets, hours matter    →  TRANSFER  →  DataSync / Transfer Family
   ALWAYS, on-prem apps keep using local protocols   →  HYBRID    →  Storage Gateway
   ONCE, far too much for the network                →  OFFLINE   →  (Snow family — see the note)

Streaming — Kinesis Data Streams

From Amazon Kinesis Data Streams terminology and concepts, verified 2026-09-25:

"A Kinesis data stream is a set of shards. Each shard has a sequence of data records."

The numbers that size a stream:

Per shard
Writes 1,000 records/s, 1 MB/s (including partition keys)
Reads 5 transactions/s, 2 MB/s
Record (data blob) up to 1 MB
Retention default 24 hours, up to 8,760 hours (365 days)

Partition key decides the shard: "An MD5 hash function is used to map partition keys to 128-bit integer values and to map associated data records to shards." Records with the same key go to the same shard, in order.

Multiple consumers read independently. "There can be multiple applications for one stream, and each application can consume data from the stream independently and concurrently." That — replay and many readers of the same records — is what separates a stream from a queue.

Capacity modes

From Choose the right mode to stream in:

number_of_shards = ceiling(max(incoming_write_bandwidth_in_KiB/1024, outgoing_read_bandwidth_in_KiB/2048))

Worked: 3,000 records/s × 2 KB = 6,000 KB/s in; two consumers → 12,000 KB/s out. max(6,000/1,024, 12,000/2,048) = max(5.86, 5.86) → 6 shards. (Records/s also caps at 1,000 per shard: 3,000 records needs at least 3, so bandwidth dominates here.)

⚠️ Hot shards. In both modes, "You may experience read and write exceptions if you are using a partition key that leads to uneven data distribution." The fix is a higher-cardinality partition key, not more shards.

Enhanced fan-out gives each consumer dedicated throughput; "Enhanced Fan-Out supports adding up to 20 consumer applications to a data stream" (on-demand standard). Use it when several consumers read the same stream and would otherwise share the 2 MB/s per shard.

Mode switches: "twice within 24 hours".

Streaming — Amazon Data Firehose

From What is Amazon Data Firehose?:

"a fully managed service for delivering real-time streaming data to destinations such as Amazon S3, Amazon Redshift, Amazon OpenSearch Service, Amazon OpenSearch Serverless, Splunk, Apache Iceberg Tables, and any custom HTTP endpoint … With Amazon Data Firehose, you don't need to write applications or manage resources."

Format conversion, from Convert input data format:

"Amazon Data Firehose can convert the format of your input data from JSON to Apache Parquet or Apache ORC before storing the data in Amazon S3. Parquet and ORC are columnar data formats that save space and enable faster queries compared to row-oriented formats like JSON."

It needs a schema in the AWS Glue Data Catalog. And ⚠️ "If you want to convert an input format other than JSON, such as comma-separated values (CSV) … you can use AWS Lambda to transform it to JSON first." So CSV → Parquet in Firehose is Lambda transform + format conversion, not conversion alone.

Data Streams vs Firehose

Kinesis Data Streams Amazon Data Firehose
You write consumers? yes (or Lambda, Firehose, etc.) no — it delivers
Replay / multiple independent readers yes, within retention no
Latency shape real time, per record buffered by size/interval
Capacity shards (provisioned) or on-demand managed
Built-in destinations none S3, Redshift, OpenSearch, Splunk, Iceberg, HTTP
JSON → Parquet no (do it downstream) yes, with Glue schema

→ "Load streaming data into S3/Redshift with the least operational overhead" ⇒ Firehose. "Multiple applications must process the same stream in real time, with replay" ⇒ Data Streams.

⚠️ The in-scope list also names Amazon MSK (managed Kafka) and Amazon Kinesis Video Streams. I did not fetch either for this lesson. A stem that says "existing Apache Kafka applications" points at MSK; read docs.aws.amazon.com/msk/latest/developerguide/what-is-msk.html before relying on specifics.

Transfer — DataSync

From What is AWS DataSync?:

"a secure, reliable, high‐speed file transfer service that helps you quickly and easily transfer your file or object data to, from, and between AWS storage services."

→ Bulk or recurring copy of files/objects into AWS over the network, with verification = DataSync. It moves data; it does not give on-premises apps a live mount (that's Storage Gateway).

Hybrid — Storage Gateway

The exam guide lists it under both 3.1 ("Hybrid storage solutions") and 3.5. From the Storage Gateway documentation: "connects an on-premises software appliance with cloud-based storage", documented by gateway type:

Gateway What on-prem sees What it's for
S3 File Gateway NFS (v3, 4.1) / SMB (2, 3) file shares files stored as S3 objects; "low-latency access to data through transparent local caching"
FSx File Gateway SMB shares "access to in-cloud Amazon FSx for Windows File Server shares from on-premises facilities"
Volume Gateway iSCSI block volumes cached or stored volumes (below)
Tape Gateway tapes "a durable, cost-effective tape-based solution for archiving data"

Volume Gateway modes, from What is Volume Gateway?:

⚠️ Cached vs stored is a latency question. "Whole dataset needs local latency" ⇒ stored. "Most data can be remote, hot subset local, shrink on-prem storage" ⇒ cached.

S3 File Gateway can also be deployed "as an AMI in Amazon EC2" or as a "hardware appliance".

⚠️ I fetched only the landing page for Tape Gateway and FSx File Gateway, not their guides. I make no claim about their current availability to new customers. Check each user guide's first page before recommending either.

Transfer — Transfer Family

From What is AWS Transfer Family?:

→ Partners already send files by SFTP and can't change = Transfer Family into S3.

⚠️ The page I fetched does not describe Transfer Family endpoint types (public vs VPC) or IP allow-listing in detail. Read docs.aws.amazon.com/transfer/latest/userguide/create-server-in-vpc.html before answering questions about restricting access to a Transfer Family endpoint.

Transfer — S3 Transfer Acceleration

From Configuring fast, secure file transfers using Amazon S3 Transfer Acceleration:

"a bucket-level feature that enables fast, easy, and secure transfers of files over long distances between your client and an S3 general purpose bucket … As the data arrives at an edge location, the data is routed to Amazon S3 over an optimized network path."

Why use it, verbatim: "Your customers upload to a centralized general purpose bucket from all over the world"; "You transfer gigabytes to terabytes of data on a regular basis across continents"; "You can't use all of your available bandwidth over the internet when uploading to Amazon S3."

Requirements that eliminate options:

AWS provides a Speed Comparison tool to test whether it helps before you pay for it.

Offline — the Snow family, in one paragraph

The in-scope services list still includes "AWS Snow Family", but no Domain 3 task statement names it. And the Snowball Edge developer guide now opens with: "AWS Snowball Edge is no longer available to new customers. New customers should explore AWS DataSync for online transfers, AWS Data Transfer Terminal for secure physical transfers, or AWS Partner solutions. For edge computing, explore AWS Outposts." The device "can transport data at speeds faster than the internet" by shipping it, with a "Storage Optimized 210 TB" configuration. ⚠️ If an exam item still offers Snowball for "petabytes, limited bandwidth, one-time migration", it is the historical answer; for new designs, AWS's own page redirects you. I did not fetch the Data Transfer Terminal documentation.

Sizes and speeds — the arithmetic

The knowledge item "Sizes and speeds needed to meet business requirements" is arithmetic the exam expects you to do in your head:

(The bandwidth figures are unit conversion, not an AWS-published number.)

Secure access to ingestion access points

What the pages I fetched actually support:

The general controls — IAM least privilege, KMS, VPC endpoint policies — are SAA1 material.

Choosing, under exam conditions

The stem says Answer
"clickstream, several apps process the same events in real time, replay for 7 days" Kinesis Data Streams (retention ≥ 7 days; enhanced fan-out)
"load IoT JSON into S3 as Parquet, no code to manage" Firehose + format conversion (Glue schema)
"CSV into S3 as Parquet via Firehose" Lambda transform to JSON, then conversion
"stream traffic unpredictable, no capacity planning" Kinesis on-demand
"one shard throttles while others idle" hot partition key → higher-cardinality key
"copy 200 TB from NFS to S3 over DX, verify integrity" DataSync
"on-prem apps keep using NFS, data lands in S3" S3 File Gateway
"on-prem iSCSI, hot data local, rest in S3" Volume Gateway cached
"partners upload via SFTP, can't change clients" Transfer Family → S3
"global users upload large files to one bucket; slow across continents" S3 Transfer Acceleration

Check yourself

  1. A provisioned stream ingests 4 MB/s and has three consumers using shared throughput. How many shards?
  2. Why can't Firehose alone turn CSV into Parquet?
  3. DataSync or Storage Gateway: on-prem applications need a live NFS mount backed by S3.
  4. Cached or stored Volume Gateway: the whole dataset needs local latency.
  5. Name two reasons Transfer Acceleration might not be available for a given bucket.
Answers
  1. Write: 4 MB/s ÷ 1 = 4. Read: 4 × 3 = 12 MB/s ÷ 2 = 6. 6 shards (or use enhanced fan-out so each consumer gets dedicated read throughput).
  2. Format conversion reads JSON only; "If you want to convert an input format other than JSON, such as comma-separated values (CSV) … you can use AWS Lambda to transform it to JSON first."
  3. Storage Gateway (S3 File Gateway) — DataSync copies data; it doesn't present a mount.
  4. Stored volumes — "If you need low-latency access to your entire dataset."
  5. The bucket name contains periods (must be DNS-compliant, no "."), or the bucket is in a Region not on the supported list. (Also: path-style requests aren't supported.)

Teaching this section

← PreviousNetwork performance — edge, hybrid links, placementNext →Transformation and analytics — Glue, EMR, Athena, Lake Formation, Quick