AWS Training
Modules Listen All tracks

← Design High-Performing Architectures

Starts this lesson and continues through 8 more to the end of certification prep.

Transformation and analytics — Glue, EMR, Athena, Lake Formation, Quick

Why this lesson closes the module

The second half of task statement 3.5. The knowledge items:

And the skills: "Building and securing data lakes", "Implementing visualization strategies", "Selecting appropriate compute options for data processing (for example, Amazon EMR)", and — the most concrete item in the whole guide — "Transforming data between formats (for example, .csv to .parquet)".

Note the guide says Amazon Quick, not QuickSight. The HTML guide picked up the rename; see ../../PLAN.md.

The canonical exam architecture is a data lake, and it has five layers:

   STORE        S3                                  (lesson 1: prefixes, multipart)
   CATALOG      AWS Glue Data Catalog  ← crawlers   (schema discovery)
   GOVERN       AWS Lake Formation                  (who sees which column / row / cell)
   TRANSFORM    Glue ETL (serverless Spark) · EMR (clusters or serverless) · Athena CTAS
   QUERY/SHOW   Athena (SQL on S3) → Amazon Quick (dashboards)

Every question in this lesson asks you to pick the right service for one layer.

Transform — AWS Glue

From What is AWS Glue?, verified 2026-09-25:

"AWS Glue is a serverless data integration service that makes it easy for analytics users to discover, prepare, move, and integrate data from multiple sources."

The pieces the exam uses:

Piece What AWS says
Data Catalog "manage your data in a centralized data catalog"; queryable "using Amazon Athena, Amazon EMR, and Amazon Redshift Spectrum"
Crawlers "automatically infer schema information and integrate it into your AWS Glue Data Catalog"
ETL jobs run on "the Apache Spark–based serverless ETL engine"; schedule, on demand, or event-triggered
Streaming ETL "Clean and transform streaming data in transit"
Glue Studio "a graphical interface … visually compose data transformation workflows"
DataBrew "a visual data preparation tool … clean and normalize data without writing any code"
FindMatches "deduplicates and finds records that are imperfect matches"

Scale and cost: "Dynamically scale resources up and down based on workload" and "pay-as-you-go billing". "It's also serverless, which means there's no infrastructure to manage."

→ "Transform data with the least operational overhead" ⇒ Glue. It's serverless Spark with a catalog attached.

Transform — Amazon EMR

From What is Amazon EMR?:

"a managed cluster platform that simplifies running big data frameworks, such as Apache Hadoop and Apache Spark, on AWS to process and analyze vast amounts of data … Amazon EMR also lets you transform and move large amounts of data into and out of other AWS data stores and databases, such as Amazon S3 and Amazon DynamoDB."

Lesson 2 covered node types (primary/core/task) and Spot on task nodes. The data-processing points:

EMR Serverless, from What is Amazon EMR Serverless?: "a deployment option for Amazon EMR that provides a serverless runtime environment" for "Apache Spark and Apache Hive". "you don't have to configure, optimize, secure, or operate clusters"; it "automatically determines the resources that the application needs … and releases the resources when the jobs finish." Pre-initialized capacity keeps "workers initialized and ready to respond in seconds".

Glue vs EMR vs EMR Serverless

Requirement Answer
serverless ETL, catalog integration, minimal ops Glue
existing Hadoop/Spark/Hive jobs, need specific frameworks, tune the cluster EMR on EC2
Spark or Hive without managing a cluster EMR Serverless
huge nightly batch, cheapest EMR transient cluster with Spot task nodes (lesson 2)

⚠️ The guide's example for "Selecting appropriate compute options for data processing" is EMR — so expect a question whose right answer is EMR even though Glue could do the job. The tell is usually an existing Hadoop/Spark codebase or a named framework.

Query — Amazon Athena

From What is Amazon Athena?:

"an interactive query service that makes it easy to analyze data directly in Amazon Simple Storage Service (Amazon S3) using standard SQL."

"Athena SQL and Apache Spark on Amazon Athena are serverless, so there is no infrastructure to set up or manage, and you pay only for the queries you run."

→ Ad-hoc SQL on data already in S3, no infrastructure ⇒ Athena. It reads table definitions from the Glue Data Catalog.

Athena performance = scan less

From Optimize Athena performance and Optimize data:

  1. Partition. "When you filter on partition key columns, only data from matching partitions is read." But: "too many partition keys can result in fragmented datasets with too many files and files that are too small." Optimise for common queries — "if your queries look at time spans of days, don't partition by hour."
  2. Partition projection computes partitions "in memory based on the query and the rules instead of looking up partitions in the AWS Glue Data Catalog" — faster for highly partitioned tables.
  3. Columnar formats. "Only the columns needed for the query are loaded", and files carry min/max metadata so pages can be skipped.
  4. Compress. "Querying compressed data is faster and also cheaper because you pay for the number of bytes scanned before decompression."
  5. Avoid many small files. "loading a single bigger file from Amazon S3 is faster than loading the same records from many smaller files." Too many can hit S3's "5,500 requests per second to a single index partition" and return "SlowDown: Please reduce your request rate" — lesson 1's prefix limit, showing up in analytics.
  6. Bucketing for "a key with high cardinality" looked up by single values.

Transform formats — CSV to Parquet

The skill "Transforming data between formats (for example, .csv to .parquet)". From Use columnar storage formats:

Parquet and ORC give, verbatim:

Three ways to convert, all named by AWS:

Method When
Athena CTAS — "CREATE TABLE AS (CTAS) queries to convert data into Parquet or ORC in one step" data already in S3 and cataloged; one-off or periodic SQL
Glue ETL job — "running jobs in AWS Glue"; "AWS Glue supports using the same technique to convert CSV data to ORC, or JSON data to either Parquet or ORC" scheduled pipelines, heavier transforms
Firehose format conversion — JSON → Parquet/ORC on the way in (record-format-conversion); CSV needs a Lambda to JSON first streaming ingestion (lesson 5)

Parquet vs ORC, per the Athena page: Parquet "might be a better choice if you plan to perform complex queries"; ORC "usually results in smaller files" and "supports a wider range of complex data types". Both support schema evolution. ⚠️ On the exam, the guide's own example is Parquet; pick Parquet unless the stem gives a reason for ORC.

Govern — AWS Lake Formation

The skill "Building and securing data lakes". From What is AWS Lake Formation?:

"AWS Lake Formation helps you centrally govern, secure, and globally share data for analytics and machine learning. With Lake Formation, you can manage fine-grained access control for your data lake data on Amazon S3 and its metadata in AWS Glue Data Catalog."

⚠️ Why S3 bucket policies aren't the answer. The unit a bucket policy talks about is an S3 object (SAA1). "Analysts see every column of this table except ssn" or "the EU team sees only EU rows" is a statement about table columns and rows — which is exactly what Lake Formation's page says its permissions model adds on top of IAM. Column-, row- or cell-level access to a data lake ⇒ Lake Formation.

Show — Amazon Quick

From What is Amazon Quick?:

"Amazon Quick is an AI-powered service for automating tasks, analyzing data, building web applications, and conducting research."

Its BI component: "Amazon Quick Sight – Interactive data visualization and business intelligence. Connect to data sources, build dashboards, and embed analytics in applications." And: "Amazon Quick evolved from Amazon QuickSight. QuickSight continues as Amazon Quick Sight, a feature within Quick. All existing QuickSight APIs, SDKs, and integrations continue to work without changes."

→ "Dashboards for business users on the data lake, serverless" ⇒ Amazon Quick (Quick Sight), usually reading via Athena.

This repo already covers Quick in depth — go there

The "Implementing visualization strategies" skill is what the quicksight/ track is for. Rather than compress it here:

If you need Read
the rename, and which doc tree to trust Q0 lesson 1 — The rename
editions, roles, identity, Regions Q0 — Platform Foundations
connecting to Athena and S3, and the two IAM roles Q1 lesson 3 — The AWS-managed sources
private connectivity to RDS/Redshift in a VPC Q1 lesson 4 — VPC connections
SPICE vs direct query, sizing and quotas Q2 — SPICE Internals and Data at Scale, especially lesson 5
the AI layer (Flows, Automate, Index, Research) Q9 — The Quick Suite AI Layer

For the exam, the depth you need is: Quick Sight is the visualization layer; it reads from Athena, Redshift, RDS, S3 and others; and it can hold an in-memory copy of data (SPICE) or query directly. The quicksight/ modules carry their own verification dates and sources — they were not re-fetched for this lesson.

Choosing, under exam conditions

The stem says Answer
"discover schemas of files landing in S3 automatically" Glue crawler → Data Catalog
"serverless ETL, least operational overhead" Glue ETL job
"existing Spark/Hadoop jobs, specific framework versions" EMR
"Spark jobs, no cluster to manage" EMR Serverless (or Glue)
"ad-hoc SQL on S3, pay per query" Athena
"Athena queries slow and expensive on CSV" convert to Parquet (CTAS or Glue), partition, compress
"thousands of tiny files, Athena slow, S3 SlowDown errors" compact into larger files
"analysts may not see the PII columns; EU team sees only EU rows" Lake Formation column permissions + data filters
"share governed tables with another account" Lake Formation cross-account sharing
"manage permissions on thousands of tables by classification" LF-Tags
"business dashboards on the lake" Amazon Quick (Quick Sight) via Athena

Check yourself

  1. Name three ways to turn CSV in S3 into Parquet, and when each fits.
  2. Why is partitioning by hour a bad idea for queries that span days?
  3. Glue or EMR: an existing Hive codebase with custom bootstrap requirements.
  4. What does Lake Formation give you that an S3 bucket policy can't?
  5. What did Amazon QuickSight become, and do its APIs still work?
Answers
  1. Athena CTAS (already in S3 and cataloged; SQL, one step), Glue ETL job (scheduled pipeline), Firehose format conversion on ingest (JSON only — CSV needs a Lambda to JSON first).
  2. "Having too many partition keys can result in fragmented datasets with too many files and files that are too small." Optimise for the common query; sort by timestamp inside day partitions instead.
  3. EMR — framework choice and bootstrap actions are what EMR gives you that Glue doesn't.
  4. Column-, row- and cell-level permissions on cataloged data, LF-Tags, cross-account sharing, and CloudTrail audit of data access — all enforced across Athena, EMR, Glue, Redshift Spectrum and Quick.
  5. Amazon Quick Sight, a feature of Amazon Quick. "All existing QuickSight APIs, SDKs, and integrations continue to work without changes."

Teaching this section

← PreviousIngestion and streaming — Kinesis, Firehose, and moving data inFinished →Cheat sheet, lab & quiz