AWS Training
Modules Listen Certification

← Data Sources and Connectivity

Starts this lesson and continues through 25 more to the end of the course.

The landscape — what connects, from where

Two lists, one service

There are two authoritative lists of what Amazon Quick Sight connects to, and they don't line up.

The user guide list (Supported data sources, verified 2026-08-16) is organized by how a human thinks: relational stores, files, SaaS, local data. The API list is the Type enum on CreateDataSource — 38 values, organized by nothing at all, and containing things the user guide page doesn't mention (GOOGLESHEETS, WEB_CRAWLER, QBUSINESS, S3_KNOWLEDGE_BASE).

You need both. The user guide tells you whether a connection is possible; the enum tells you what to type. When you automate — and Q7 assumes you will — the enum is the one that counts.

The relational sources

From the user guide, verbatim list (verified 2026-08-16): Amazon Athena, Amazon Aurora, AWS Glue Data Catalog ("accessed using AWS Glue data catalog compatible services, such as Athena or Redshift Spectrum"), Amazon OpenSearch Service, Amazon Redshift, Amazon Redshift Spectrum, Amazon S3, Amazon S3 Analytics, Amazon S3 Tables, Apache Impala, Apache Spark 2.0+, AWS IoT Analytics, Databricks (E2 platform only, Spark 1.6–3.0), Exasol 7.1.2+, Google BigQuery, MariaDB 10.0+, Microsoft SQL Server 2012+, MySQL 5.7+, Oracle 12c+, PostgreSQL 9.3.1+, Presto 0.167+, Snowflake, Starburst, Trino, Teradata 14.0+, Timestream.

Notice what's not a data source: Glue is not directly connectable. You reach Glue Data Catalog tables through Athena or Redshift Spectrum. Teams migrating from other BI tools ask for a "Glue connector" constantly; the answer is "point Athena at it."

Three version notes from the same page that turn into tickets:

Where the database is allowed to live

This paragraph of the user guide does more work than any other (verbatim, verified 2026-08-16):

"Amazon Redshift clusters, Amazon Athena databases, and Amazon RDS instances must be in AWS. Other database instances must be in one of the following environments to be accessible from Amazon Quick Sight: Amazon EC2, Local (on-premises) databases, Data in a data center or some other internet-accessible environment."

So a self-managed PostgreSQL box in your basement is a legitimate data source — if there's a network path (lesson 4). But you cannot point Quick Sight at, say, a Redshift-wire-compatible service running outside AWS and call it Redshift.

Files

Supported formats: CSV/TSV, ELF/CLF (extended/common log), JSON, XLSX. Encoding is UTF-8 — and explicitly not UTF-8 with BOM. Files in S3 compressed with zip or gzip import as-is; any other compression, or compressed files uploaded from your local network, must be decompressed first.

The JSON support has documented edges (all verbatim from the supported-data-sources page): schema and type inference, flattening, and embedded-object parsing are supported; not supported are "a structure containing a list of records", list attributes ("skipped during import"), custom upload settings, and — the quiet one — "error messaging for invalid JSON". A malformed JSON file doesn't necessarily announce itself. Remember RowsDropped from Q2 lesson 3? This is one of the things that feeds it.

Per-file shape limits, from Data source quotas (verified 2026-08-16):

Limit Value
Columns per file (or per query result set) 2,000
Characters per column name 127
Characters per field 2,047 (65,534 in the "new data preparation experience")
Files per S3 manifest 1,000

⚠️ The 2,000-column limit doesn't block the import — "File imports and query result sets can contain more than 2,000 columns", but you must then manually exclude fields in dataset settings until fewer than 2,000 remain (per the Service Quotas description of "Data Prep: Fields per dataset", which is not adjustable). Wide event tables hit this.

SaaS sources

Direct connection: Jira, ServiceNow. Via OAuth (authorize on the SaaS site): Adobe Analytics, GitHub, Salesforce. For Salesforce, only Enterprise, Unlimited, and Developer editions are supported as sources.

Hold onto the Jira/ServiceNow pair — lesson 2 shows they're also the two sources that can't use Secrets Manager credentials.

The API enum — and its fossils

The full Type valid values, verbatim from CreateDataSource (verified 2026-08-16):

ADOBE_ANALYTICS | AMAZON_ELASTICSEARCH | ATHENA | AURORA | AURORA_POSTGRESQL |
AWS_IOT_ANALYTICS | GITHUB | JIRA | MARIADB | MYSQL | ORACLE | POSTGRESQL |
PRESTO | REDSHIFT | S3 | S3_TABLES | SALESFORCE | SERVICENOW | SNOWFLAKE |
SPARK | SQLSERVER | TERADATA | TWITTER | TIMESTREAM | AMAZON_OPENSEARCH |
EXASOL | DATABRICKS | STARBURST | TRINO | BIGQUERY | GOOGLESHEETS |
GOOGLE_DRIVE | CONFLUENCE | SHAREPOINT | ONE_DRIVE | WEB_CRAWLER |
S3_KNOWLEDGE_BASE | QBUSINESS

Thirty-eight values, and they are a geological record:

Check yourself

  1. A team wants to connect Quick Sight "directly to Glue." What do you tell them?
  2. Your compliance team mandates TLS 1.2 for the Aurora MySQL connection. What do you check, and where?
  3. Which two facts make an S3-hosted .csv.gz file importable as-is, and what breaks if the same file is bzip2-compressed?
  4. You're writing Terraform for an OpenSearch data source. Which Type does AWS document you should use, and what's the trap?
  5. A 2,300-column query result imports successfully. What must happen before the dataset is usable, and is the ceiling adjustable?
Answers
  1. There is no Glue connector. Glue Data Catalog is reached through Athena or Redshift Spectrum — create an Athena data source pointed at the catalog.
  2. The database engine version. TLS 1.2 requires MySQL 5.7.28+; below that Quick Sight falls back to TLS 1.1 with no error. The control lives in RDS, not in Quick Sight settings.
  3. S3 origin + gzip (or zip) compression — those import as-is. bzip2 isn't on the list, so you'd decompress before import.
  4. AMAZON_ELASTICSEARCH — the CreateDataSource page says to use it for OpenSearch Service. The trap is that AMAZON_OPENSEARCH also exists and is also valid, so two data sources pointing at the same domain can carry different types.
  5. Fields must be manually excluded in dataset settings until fewer than 2,000 remain. "Data Prep: Fields per dataset = 2,000" is marked not adjustable in the Service Quotas table.

Teaching this section

Next →Creating data sources through the API