AWS Training
Modules Listen Certification

← Data Sources and Connectivity

Starts this lesson and continues through 21 more to the end of the course.

The failures that look like network failures

Two error surfaces, and people read the wrong one

A "connection problem" in Quick Sight reports itself in two different places, depending on when it fails:

The first enum is small and blunt; the second is large and precise. When a refresh fails, go to the ingestion error — don't re-test the data source first. When creating fails, the ingestion history has nothing for you. Knowing which surface you're on is half the diagnosis.

The data source error enum

From DataSourceErrorInfo (verified 2026-08-16), Type valid values, verbatim:

ACCESS_DENIED | COPY_SOURCE_NOT_FOUND | TIMEOUT | ENGINE_VERSION_NOT_SUPPORTED |
UNKNOWN_HOST | GENERIC_SQL_FAILURE | CONFLICT | UNKNOWN

Mapped onto the module's triage order (the mapping is mine — the enum values are documented, the per-value interpretations follow from their names, so quote the value, not my gloss, to Support):

Triage stage Enum value First move
1 — NAME UNKNOWN_HOST DNS: typo in Host, private zone without DnsResolvers, wrong Region
2 — ROUTE TIMEOUT SGs (the three rules), subnets, AvailabilityStatus, DB actually up?
3 — AUTH ACCESS_DENIED credentials/secret, the two service roles, KMS grants
4 — QUERY GENERIC_SQL_FAILURE the SQL, the schema, the driver
— ENGINE_VERSION_NOT_SUPPORTED lesson 1's version floors (MySQL 5.7+, Postgres 9.3.1+, …)
— COPY_SOURCE_NOT_FOUND the lesson-2 CopySourceArn chain broke — the source was deleted
— CONFLICT / UNKNOWN state collision / no signal — evidence-gathering mode

The crucial split is the first three rows. UNKNOWN_HOST and TIMEOUT are connectivity; ACCESS_DENIED is not, no matter how network-shaped the symptom feels. The name → route → auth ladder exists because each stage's failure masks the stages below it: you cannot see an auth problem through an unresolvable hostname.

The Q2 ingestion enum makes the same split with sharper teeth — recall UNRESOLVABLE_HOST (DNS) vs UNROUTABLE_HOST (resolved, but no path). The ingestion enum literally distinguishes triage stages 1 and 2 for you. Use it.

The timeout that isn't a network timeout

From Data source quotas (verified 2026-08-16): for direct query, "Amazon Quick Sight generates a timeout after 2 minutes" when generating a visual (the Service Quotas table states it as "Query timeout for visuals: 120 seconds", not adjustable; dataset previews get 45 seconds).

And the sentence that explains a thousand mystery incidents, verbatim:

"However, not all database drivers react to the 2-minute timeout, for example Amazon Redshift. In these cases, the query runs for as long as it takes for the response to return, which can result in long-running queries on your database."

Put those together: the visual dies at 120s and looks like a connection failure — while the query is still running on Redshift. The user retries. Another orphan query starts. Twenty minutes of "Quick Sight can't connect to Redshift" later, the cluster is grinding under a pile of zombie queries, which slows every new query past 120s, which "confirms" the connection problem. The docs' own advice: cancel the queries from the database side (they link Redshift's cancel_query procedure), then either optimize the SQL or move the dataset to SPICE.

⚠️ When "Quick Sight can't connect" coincides with "the database is slow", check the database's running-queries view first. The BI tool is often the source of the load, not the victim of it.

The triage runbook

0. WHICH SURFACE?   create/test failed → DescribeDataSource.ErrorInfo
                    refresh failed     → ListIngestions → ErrorInfo (Q2.3)
                    visual slow/dead   → suspect the 120s timeout; check DB load

1. NAME    UNKNOWN_HOST / UNRESOLVABLE_HOST
           → dig the hostname from a VPC instance; check DnsResolvers; check Region
2. ROUTE   TIMEOUT / UNROUTABLE_HOST
           → describe-vpc-connection (AvailabilityStatus!); the three SG rules;
             is the DB accepting connections at all?
3. AUTH    ACCESS_DENIED / PERMISSION_DENIED / auth-family
           → which service role exists (lesson 3)? secret intact (lesson 2 — did
             someone touch the console)? KMS grant? results bucket?
4. QUERY   GENERIC_SQL_FAILURE / SQL_* family
           → run the same SQL as the same principal directly against the source

Evidence for a support case, same discipline as Q2 lesson 5: account ID, Region, data source ID and ARN, the DataSourceErrorInfo.Type enum value verbatim, the RequestId of the failing call, timestamps, and — for VPC cases — the VPC connection ID and its AvailabilityStatus. A case that opens with "TIMEOUT on data source X, VPC connection Y is PARTIALLY_AVAILABLE, RequestId Z" skips the entire first week of back-and-forth.

Check yourself

  1. A refresh fails on an existing Redshift data source. Which error surface do you read, and why is it the better one?
  2. ACCESS_DENIED after infrastructure "didn't change" — name three lesson-specific causes.
  3. Users report intermittent "can't connect to Redshift" and the DBA reports the cluster is slow. Describe the loop that's probably running, and the two-step fix.
  4. Why does the name → route → auth order matter — why not check auth first?
  5. COPY_SOURCE_NOT_FOUND — what happened, and which lesson-2 decision would have prevented it?
Answers
  1. The ingestion's ErrorInfo via ListIngestions — 45 precise values (including the DNS-vs-routing split) against the data source enum's 8 blunt ones.
  2. A secret silently dropped by a console edit; a fix applied to aws-quicksight-service-role-v0 while aws-quicksight-s3-consumers-role-v0 exists and is being used; a KMS key rotation/grant missing. (Also fair: results-bucket access revoked by someone else's checkbox change.)
  3. Visuals time out at 120s but Redshift's driver ignores the timeout, so each retry stacks another still-running query on the cluster; the load slows new queries past 120s, reinforcing the "connection" diagnosis. Fix: cancel the queries on the Redshift side, then optimize the SQL or move to SPICE.
  4. Each stage masks the ones below: an unresolvable name prevents any routing evidence; a dead route prevents any auth evidence. Checking auth first burns time on a layer that may never have been reached.
  5. The data source borrowed credentials via CopySourceArn and the source data source was deleted. Using a SecretArn (credentials owned by Secrets Manager, not another data source) avoids the chain entirely.

Teaching this section

← PreviousVPC connections — how the packets actually moveFinished →Cheat sheet, lab & quiz