AWS Training
Modules Listen Certification

← SPICE Internals and Data at Scale

Starts this lesson and continues through 18 more to the end of the course.

Ingestion mechanics and the failure taxonomy

The console tells you almost nothing

The refresh history in the Quick Sight console gives you a red "Failed" and a short message. That is not enough to open a support case with, and it is not enough to fix most problems.

The API gives you nine fields per ingestion, including a 45-value error type enum, the number of rows that were silently dropped, and whether the run was manual or scheduled. Everything in this lesson is about reading those nine fields, because that is what turns "the refresh failed" into a diagnosis.

The ingestion object

From Ingestion (verified 2026-08-09):

Field Type Why you care
Arn String Identity — goes in the support case
IngestionId String Your handle for this run; you choose it on create
CreatedTime Timestamp When it started
IngestionStatus Enum INITIALIZED | QUEUED | RUNNING | FAILED | COMPLETED | CANCELLED
RequestSource Enum MANUAL | SCHEDULED
RequestType Enum INITIAL_INGESTION | EDIT | INCREMENTAL_REFRESH | FULL_REFRESH
IngestionSizeInBytes Long What actually landed
IngestionTimeInSeconds Long Duration — your trend line for "is this getting worse?"
RowInfo Object RowsIngested, RowsDropped, TotalRowsInDataset
QueueInfo Object QueuedIngestion, WaitingOnIngestion
ErrorInfo Object Type (the 45-value enum) and Message

RowsDropped is the field nobody looks at

From RowInfo:

An ingestion can report COMPLETED with a non-zero RowsDropped. The console shows you a green tick. The dashboard shows numbers that are quietly wrong.

⚠️ This is the most dangerous failure mode in the whole service — not the refresh that fails loudly, but the one that succeeds while dropping 2% of your rows. Nobody notices for a quarter, and then somebody reconciles against the source system.

Alert on RowsDropped > 0, not just on IngestionStatus == FAILED. If you take one operational practice out of this module, take that one.

RequestType distinguishes four different things

INITIAL_INGESTION is the first load. FULL_REFRESH and INCREMENTAL_REFRESH are what you'd expect. EDIT is the one people forget: saving a dataset triggers an ingestion. An author tweaking a calculated field at 4pm can start a full reload of a two-billion-row dataset without realising it.

QueueInfo explains "why is nothing happening"

QueuedIngestion and WaitingOnIngestion tell you this run is blocked behind another one on the same dataset. A status of QUEUED with a populated QueueInfo is not a failure — it's a concurrency answer, and it's the correct response to "the refresh has been running for an hour" when in fact it hasn't started.

Starting an ingestion

CreateIngestion (verified 2026-08-09):

PUT /accounts/{AwsAccountId}/data-sets/{DataSetId}/ingestions/{IngestionId} HTTP/1.1
Content-type: application/json

{
   "IngestionType": "string"
}

IngestionType is optional, and its valid values are INCREMENTAL_REFRESH | FULL_REFRESH. IngestionId is yours to choose: 1–128 characters matching ^[a-zA-Z0-9-_]+$.

Via the CLI:

aws quicksight create-ingestion \
  --aws-account-id 111122223333 \
  --data-set-id sales-fact-v3 \
  --ingestion-id "manual-$(date +%Y%m%dT%H%M%S)" \
  --ingestion-type FULL_REFRESH \
  --region us-east-1

Response: Arn, IngestionId, IngestionStatus, RequestId.

Put a timestamp in the IngestionId. It's a PUT to a path containing the ID, and ResourceExistsException (HTTP 409) is a documented error. A fixed ID works exactly once.

⚠️ Keep the RequestId from every failed call. AWS Support can look up a request by its ID. A case that quotes one gets a different quality of answer to a case that describes a symptom.

Documented errors on CreateIngestion

Error HTTP What it usually means
AccessDeniedException 401 Credentials, or the identity isn't authorised for Quick Sight
InvalidParameterValueException 400 Bad IngestionType, malformed ID
LimitExceededException 409 You've hit a limit — including the 32-call quota
ResourceExistsException 409 That IngestionId already exists
ResourceNotFoundException 404 Wrong account, dataset, or Region
ThrottlingException 429 Rate limiting — back off and retry
InternalFailureException 500 AWS side. Retry, then open a case with the RequestId

Note that LimitExceededException returns 409, not 429. A retry loop that only backs off on 429 will hammer a 409 forever without ever succeeding — because the 32-call quota won't clear for hours. Treat 409 LimitExceededException as terminal for the day, not as a retryable error.

Reading ingestion history

ListIngestions is "Limited to 5 TPS per user and 25 TPS per account", MaxResults 1–100.

aws quicksight list-ingestions \
  --aws-account-id 111122223333 \
  --data-set-id sales-fact-v3 \
  --max-results 100 \
  --region us-east-1 \
  --query 'Ingestions[].{
      When:CreatedTime,
      Status:IngestionStatus,
      Src:RequestSource,
      Type:RequestType,
      Secs:IngestionTimeInSeconds,
      Bytes:IngestionSizeInBytes,
      In:RowInfo.RowsIngested,
      Dropped:RowInfo.RowsDropped,
      Err:ErrorInfo.Type
    }' --output table

That one command is your entire first-line diagnosis. Run it before you do anything else.

Settling the "do scheduled refreshes count?" question

Lesson 2 flagged that the docs don't say whether scheduled refreshes consume the 32-call CreateIngestion quota. RequestSource lets you answer it for your own account:

# How many ingestions in the last 24h, split by source?
aws quicksight list-ingestions \
  --aws-account-id 111122223333 --data-set-id sales-fact-v3 \
  --max-results 100 --region us-east-1 \
  --query 'Ingestions[].RequestSource' --output text | tr '\t' '\n' | sort | uniq -c

Run a controlled test: note the count, let scheduled refreshes run, then manually call create-ingestion repeatedly until you get LimitExceededException, and see whether the number of manual calls you got was 32 or 32-minus-the-scheduled-ones. That is an afternoon of work and it converts an assumption into a measurement.

Do this before you design around it, not after. And it's a perfectly good thing to ask AWS Support to confirm in writing.

The failure taxonomy — 45 error types

The complete ErrorInfo.Type enum, verbatim from ErrorInfo (verified 2026-08-09).

A note on what follows. AWS publishes the enum values but not a per-value description. The grouping below, and the "what it points at" column, are my categorisation from the value names, not AWS documentation. Use it as a triage aid. Do not quote it to AWS Support as though it were documented behaviour — quote the enum value, which is documented.

Group 1 — Capacity and size (you hit a ceiling)

Type Points at
DATA_SET_SIZE_LIMIT_EXCEEDED Per-dataset quota — not adjustable, redesign
ROW_SIZE_LIMIT_EXCEEDED An individual row too large
ACCOUNT_CAPACITY_LIMIT_EXCEEDED Account SPICE capacity — buy more or free some
SOURCE_RESOURCE_LIMIT_EXCEEDED The source ran out, not Quick Sight
SOURCE_API_LIMIT_EXCEEDED_FAILURE Source API throttling

The first three are the ones to get right: ACCOUNT_CAPACITY_LIMIT_EXCEEDED is a purchase, DATA_SET_SIZE_LIMIT_EXCEEDED is a redesign. Same red banner in the console, completely different response.

Group 2 — Identity and permission

Type Points at
FAILURE_TO_ASSUME_ROLE The Quick Sight service role can't be assumed — trust policy
IAM_ROLE_NOT_AVAILABLE Role missing or deleted
PERMISSION_DENIED Authorisation on the resource
PERMISSION_NOT_FOUND Expected permission object absent
PASSWORD_AUTHENTICATION_FAILURE Stored credentials wrong or rotated
OAUTH_TOKEN_FAILURE OAuth token expired or revoked
DATA_SOURCE_AUTH_FAILED Auth at the data source

⚠️ Credential rotation is the classic 3am cause. Nothing changed in Quick Sight; a secret rotated somewhere else and the next scheduled refresh died.

Group 3 — Network and connectivity

Type Points at
CONNECTION_FAILURE Generic connection failure
DATA_SOURCE_CONNECTION_FAILED Connection to the configured source
UNRESOLVABLE_HOST DNS — the name didn't resolve
UNROUTABLE_HOST Routing — resolved, no path. Security group / route table / VPC connection
SSL_CERTIFICATE_VALIDATION_FAILURE TLS trust chain

UNRESOLVABLE_HOST versus UNROUTABLE_HOST is a genuinely useful split: the first sends you to Route 53 or a private hosted zone, the second to security groups, NACLs, and the VPC connection. Distinguishing them saves an hour.

Group 4 — Source data and schema

Type Points at
SQL_SCHEMA_MISMATCH_ERROR Source schema changed under the dataset
SQL_TABLE_NOT_FOUND Table renamed or dropped
SQL_EXCEPTION Generic SQL error
SQL_INVALID_PARAMETER_VALUE Bad parameter in custom SQL
SQL_NUMERIC_OVERFLOW Value out of range for its target type
QUERY_TIMEOUT The source query didn't finish
INVALID_DATE_FORMAT Date parsing
DUPLICATE_COLUMN_NAMES_FOUND Two columns with the same name — common after a join
INVALID_DATAPREP_SYNTAX A calculated field or prep step is invalid
DATA_TOLERANCE_EXCEPTION Too many bad rows — related to RowsDropped
CURSOR_NOT_ENABLED / ELASTICSEARCH_CURSOR_NOT_ENABLED Source-side cursor setting

⚠️ SQL_SCHEMA_MISMATCH_ERROR has an explicit remedy in the user guide, and it is not intuitive:

"If there is a schema change in a database, Quick Sight will not be able to auto-detect it, resulting in an ingestion failure. Edit and save the dataset to update the schema and avoid ingestion failures." — Refreshing SPICE data

Quick Sight does not pick up source schema changes on its own. Someone has to open the dataset and save it. That is a hard dependency between your data platform's deploy process and your BI layer, and it belongs in your change checklist.

Group 5 — Files and objects

Type Points at
S3_FILE_INACCESSIBLE Bucket policy, KMS key, or the object moved
S3_UPLOADED_FILE_DELETED The uploaded file backing the dataset is gone
S3_MANIFEST_ERROR Malformed manifest, or over the 1,000-file limit
FAILURE_TO_PROCESS_JSON_FILE JSON couldn't be parsed
IOT_FILE_NOT_FOUND / IOT_DATA_SET_FILE_EMPTY IoT Analytics source

Group 6 — Lifecycle and state (mostly not bugs)

Type Points at
INGESTION_SUPERSEDED A newer ingestion replaced this one
INGESTION_CANCELED Cancelled
REFRESH_SUPPRESSED_BY_EDIT Someone was editing the dataset
DATA_SET_DELETED Dataset deleted mid-flight
DATA_SET_NOT_SPICE Dataset is direct query — there is nothing to ingest
SPICE_TABLE_NOT_FOUND Internal SPICE table missing
DATA_SOURCE_NOT_FOUND Data source deleted

INGESTION_SUPERSEDED and REFRESH_SUPPRESSED_BY_EDIT are the two that get wrongly escalated. Neither is a defect. They mean "two things were happening at once and this one lost". If your alerting pages someone for these, fix the alerting.

Group 7 — Everything else

Type Points at
CUSTOMER_ERROR Generic "your side"
INVALID_DATA_SOURCE_CONFIG Data source configuration
INTERNAL_SERVICE_ERROR AWS side — this is a support case

INTERNAL_SERVICE_ERROR is the one type where the correct first action is to open a case rather than investigate. Retry once; if it recurs, escalate with the ingestion ARN, IngestionId, RequestId, Region, and timestamps.

Check yourself

  1. An ingestion is COMPLETED and the dashboard totals are 3% low. Which field do you read?
  2. create-ingestion returns HTTP 409. What are the two possibilities and how do you tell them apart?
  3. What does RequestType: EDIT tell you about who caused a full reload?
  4. You get UNROUTABLE_HOST. Name two things you'd check, and one you wouldn't.
  5. The source team added a column last night and this morning's refresh failed. What's the error type, and what's the fix?
Answers
  1. RowInfo.RowsDropped. A completed ingestion can silently drop rows; alert on this, not just on FAILED.
  2. ResourceExistsException (the IngestionId was reused) or LimitExceededException (a quota, likely the 32-call one). The exception name in the response body distinguishes them. Only the first is fixed by retrying with a new ID.
  3. Somebody opened the dataset and saved it. Saving a dataset triggers an ingestion — often a full reload nobody intended.
  4. Check: security groups, route tables / VPC connection configuration, NACLs. Don't check DNS — UNROUTABLE_HOST means the name resolved; DNS problems surface as UNRESOLVABLE_HOST.
  5. SQL_SCHEMA_MISMATCH_ERROR. Quick Sight cannot auto-detect source schema changes. Someone must edit and save the dataset to refresh its schema.

Teaching this section

← PreviousThe quotas that stop youNext →Refresh strategy at scale