Pair your devices with a code and playback position follows you: pause on this device, hit resume on the other. Position is saved to the site every minute and on pause.
Open this panel on your other device and enter the same code.
Starts this lesson and continues through 8 more to the end of certification prep.
The second half of task statement 3.5. The knowledge items:
And the skills: "Building and securing data lakes", "Implementing visualization strategies", "Selecting appropriate compute options for data processing (for example, Amazon EMR)", and — the most concrete item in the whole guide — "Transforming data between formats (for example, .csv to .parquet)".
Note the guide says Amazon Quick, not QuickSight. The HTML guide picked up the rename; see
../../PLAN.md.
The canonical exam architecture is a data lake, and it has five layers:
STORE S3 (lesson 1: prefixes, multipart)
CATALOG AWS Glue Data Catalog ← crawlers (schema discovery)
GOVERN AWS Lake Formation (who sees which column / row / cell)
TRANSFORM Glue ETL (serverless Spark) · EMR (clusters or serverless) · Athena CTAS
QUERY/SHOW Athena (SQL on S3) → Amazon Quick (dashboards)
Every question in this lesson asks you to pick the right service for one layer.
From What is AWS Glue?, verified 2026-09-25:
"AWS Glue is a serverless data integration service that makes it easy for analytics users to discover, prepare, move, and integrate data from multiple sources."
The pieces the exam uses:
| Piece | What AWS says |
|---|---|
| Data Catalog | "manage your data in a centralized data catalog"; queryable "using Amazon Athena, Amazon EMR, and Amazon Redshift Spectrum" |
| Crawlers | "automatically infer schema information and integrate it into your AWS Glue Data Catalog" |
| ETL jobs | run on "the Apache Spark–based serverless ETL engine"; schedule, on demand, or event-triggered |
| Streaming ETL | "Clean and transform streaming data in transit" |
| Glue Studio | "a graphical interface … visually compose data transformation workflows" |
| DataBrew | "a visual data preparation tool … clean and normalize data without writing any code" |
| FindMatches | "deduplicates and finds records that are imperfect matches" |
Scale and cost: "Dynamically scale resources up and down based on workload" and "pay-as-you-go billing". "It's also serverless, which means there's no infrastructure to manage."
→ "Transform data with the least operational overhead" ⇒ Glue. It's serverless Spark with a catalog attached.
From What is Amazon EMR?:
"a managed cluster platform that simplifies running big data frameworks, such as Apache Hadoop and Apache Spark, on AWS to process and analyze vast amounts of data … Amazon EMR also lets you transform and move large amounts of data into and out of other AWS data stores and databases, such as Amazon S3 and Amazon DynamoDB."
Lesson 2 covered node types (primary/core/task) and Spot on task nodes. The data-processing points:
EMR Serverless, from What is Amazon EMR Serverless?: "a deployment option for Amazon EMR that provides a serverless runtime environment" for "Apache Spark and Apache Hive". "you don't have to configure, optimize, secure, or operate clusters"; it "automatically determines the resources that the application needs … and releases the resources when the jobs finish." Pre-initialized capacity keeps "workers initialized and ready to respond in seconds".
| Requirement | Answer |
|---|---|
| serverless ETL, catalog integration, minimal ops | Glue |
| existing Hadoop/Spark/Hive jobs, need specific frameworks, tune the cluster | EMR on EC2 |
| Spark or Hive without managing a cluster | EMR Serverless |
| huge nightly batch, cheapest | EMR transient cluster with Spot task nodes (lesson 2) |
⚠️ The guide's example for "Selecting appropriate compute options for data processing" is EMR — so expect a question whose right answer is EMR even though Glue could do the job. The tell is usually an existing Hadoop/Spark codebase or a named framework.
From What is Amazon Athena?:
"an interactive query service that makes it easy to analyze data directly in Amazon Simple Storage Service (Amazon S3) using standard SQL."
"Athena SQL and Apache Spark on Amazon Athena are serverless, so there is no infrastructure to set up or manage, and you pay only for the queries you run."
→ Ad-hoc SQL on data already in S3, no infrastructure ⇒ Athena. It reads table definitions from the Glue Data Catalog.
From Optimize Athena performance and Optimize data:
The skill "Transforming data between formats (for example, .csv to .parquet)". From Use columnar storage formats:
Parquet and ORC give, verbatim:
Three ways to convert, all named by AWS:
| Method | When |
|---|---|
Athena CTAS — "CREATE TABLE AS (CTAS) queries to convert data into Parquet or ORC in one step" |
data already in S3 and cataloged; one-off or periodic SQL |
| Glue ETL job — "running jobs in AWS Glue"; "AWS Glue supports using the same technique to convert CSV data to ORC, or JSON data to either Parquet or ORC" | scheduled pipelines, heavier transforms |
| Firehose format conversion — JSON → Parquet/ORC on the way in (record-format-conversion); CSV needs a Lambda to JSON first | streaming ingestion (lesson 5) |
Parquet vs ORC, per the Athena page: Parquet "might be a better choice if you plan to perform complex queries"; ORC "usually results in smaller files" and "supports a wider range of complex data types". Both support schema evolution. ⚠️ On the exam, the guide's own example is Parquet; pick Parquet unless the stem gives a reason for ORC.
The skill "Building and securing data lakes". From What is AWS Lake Formation?:
"AWS Lake Formation helps you centrally govern, secure, and globally share data for analytics and machine learning. With Lake Formation, you can manage fine-grained access control for your data lake data on Amazon S3 and its metadata in AWS Glue Data Catalog."
⚠️ Why S3 bucket policies aren't the answer. The unit a bucket policy talks about is an S3 object
(SAA1). "Analysts see every column of this table except ssn" or "the EU team sees only EU rows" is a
statement about table columns and rows — which is exactly what Lake Formation's page says its
permissions model adds on top of IAM. Column-, row- or cell-level access to a data lake ⇒ Lake
Formation.
From What is Amazon Quick?:
"Amazon Quick is an AI-powered service for automating tasks, analyzing data, building web applications, and conducting research."
Its BI component: "Amazon Quick Sight – Interactive data visualization and business intelligence. Connect to data sources, build dashboards, and embed analytics in applications." And: "Amazon Quick evolved from Amazon QuickSight. QuickSight continues as Amazon Quick Sight, a feature within Quick. All existing QuickSight APIs, SDKs, and integrations continue to work without changes."
→ "Dashboards for business users on the data lake, serverless" ⇒ Amazon Quick (Quick Sight), usually reading via Athena.
The "Implementing visualization strategies" skill is what the quicksight/ track is for. Rather than
compress it here:
| If you need | Read |
|---|---|
| the rename, and which doc tree to trust | Q0 lesson 1 — The rename |
| editions, roles, identity, Regions | Q0 — Platform Foundations |
| connecting to Athena and S3, and the two IAM roles | Q1 lesson 3 — The AWS-managed sources |
| private connectivity to RDS/Redshift in a VPC | Q1 lesson 4 — VPC connections |
| SPICE vs direct query, sizing and quotas | Q2 — SPICE Internals and Data at Scale, especially lesson 5 |
| the AI layer (Flows, Automate, Index, Research) | Q9 — The Quick Suite AI Layer |
For the exam, the depth you need is: Quick Sight is the visualization layer; it reads from Athena,
Redshift, RDS, S3 and others; and it can hold an in-memory copy of data (SPICE) or query directly. The
quicksight/ modules carry their own verification dates and sources — they were not re-fetched for this
lesson.
| The stem says | Answer |
|---|---|
| "discover schemas of files landing in S3 automatically" | Glue crawler → Data Catalog |
| "serverless ETL, least operational overhead" | Glue ETL job |
| "existing Spark/Hadoop jobs, specific framework versions" | EMR |
| "Spark jobs, no cluster to manage" | EMR Serverless (or Glue) |
| "ad-hoc SQL on S3, pay per query" | Athena |
| "Athena queries slow and expensive on CSV" | convert to Parquet (CTAS or Glue), partition, compress |
| "thousands of tiny files, Athena slow, S3 SlowDown errors" | compact into larger files |
| "analysts may not see the PII columns; EU team sees only EU rows" | Lake Formation column permissions + data filters |
| "share governed tables with another account" | Lake Formation cross-account sharing |
| "manage permissions on thousands of tables by classification" | LF-Tags |
| "business dashboards on the lake" | Amazon Quick (Quick Sight) via Athena |
quicksight/ track rather than expanding this lesson.
For the exam, knowing it's the visualization layer is enough.