AWS Training
Modules Listen Certification

← SPICE Internals and Data at Scale

Runbook — SPICE ingestion failures

For someone under pressure. No background. Teaching lives in the lessons; this page is symptom → command → decision.

export ACCOUNT=<12-digit account id>
export REGION=<the Quick Sight account's region>
export DATASET=<dataset id>

0. The one command

aws quicksight list-ingestions --aws-account-id $ACCOUNT --data-set-id $DATASET \
  --region $REGION --max-results 20 \
  --query 'Ingestions[].{When:CreatedTime,Status:IngestionStatus,Src:RequestSource,
      Type:RequestType,Secs:IngestionTimeInSeconds,In:RowInfo.RowsIngested,
      Dropped:RowInfo.RowsDropped,Err:ErrorInfo.Type}' --output table

Wrong Region ⇒ ResourceNotFoundException. Check the Region before you check anything else.


Symptom: dashboard is stale

IngestionStatus Meaning Do this
COMPLETED, RowsDropped = 0 Refresh worked Problem is elsewhere — filters, permissions, or the source itself was stale
COMPLETED, RowsDropped > 0 Silently wrong data Go to Dropped rows below
QUEUED + QueueInfo populated Blocked behind another run Wait. Not an incident.
RUNNING far past normal IngestionTimeInSeconds Slow source or capacity pressure Check source load; check SPICE capacity
FAILED Read ErrorInfo.Type Table below
No recent rows at all The schedule is gone Check schedules — an automation may have deleted and not recreated

Symptom: FAILED — route by ErrorInfo.Type

Type First action Escalate to
ACCOUNT_CAPACITY_LIMIT_EXCEEDED Admin → SPICE capacity. Purchase, or release unused after removing datasets Billing owner
DATA_SET_SIZE_LIMIT_EXCEEDED Stop. Per-dataset quota, not adjustable Data engineering — redesign
ROW_SIZE_LIMIT_EXCEEDED Find the oversized row/column at source Data engineering
SQL_SCHEMA_MISMATCH_ERROR Open the dataset, save it, refresh Whoever changed the source schema
SQL_TABLE_NOT_FOUND Table renamed/dropped at source Source owner
FAILURE_TO_ASSUME_ROLE / IAM_ROLE_NOT_AVAILABLE Check the role and its trust policy Cloud platform
PASSWORD_AUTHENTICATION_FAILURE / OAUTH_TOKEN_FAILURE / DATA_SOURCE_AUTH_FAILED Credential rotated. Update the data source Secrets owner
UNRESOLVABLE_HOST DNS. Route 53 / private hosted zone Networking
UNROUTABLE_HOST Routing. Security groups, NACLs, VPC connection Networking
SSL_CERTIFICATE_VALIDATION_FAILURE Cert chain at the source Source owner
QUERY_TIMEOUT Source query too slow. Simplify, or pre-aggregate Data engineering
S3_FILE_INACCESSIBLE / S3_MANIFEST_ERROR Bucket policy, KMS key, object moved, or > 1,000 files in the manifest Data engineering
DUPLICATE_COLUMN_NAMES_FOUND Two same-named columns, usually after a join. Alias them Dataset owner
INGESTION_SUPERSEDED / REFRESH_SUPPRESSED_BY_EDIT Not a defect. Two operations overlapped Nobody. Fix the alert.
DATA_SET_NOT_SPICE Dataset is direct query; nothing to ingest Fix the automation
INTERNAL_SERVICE_ERROR Retry once. If it recurs → support case AWS Support

Full enum: ErrorInfo


Symptom: dropped rows on a COMPLETED ingestion

  1. Record RowsIngested, RowsDropped, TotalRowsInDataset and the timestamp.
  2. Look for DATA_TOLERANCE_EXCEPTION, INVALID_DATE_FORMAT, SQL_NUMERIC_OVERFLOW in nearby runs — the same bad data usually shows up as an error type before or after.
  3. Tell the data consumers immediately. A dashboard reporting confidently wrong numbers is worse than a broken one, and someone is currently making a decision on it.
  4. Then find the offending rows at source.

Prevention: alert on RowsDropped > 0, not only on IngestionStatus == FAILED.


Symptom: create-ingestion fails

HTTP Exception Action
409 LimitExceededException Quota, probably the 32/24h. Terminal for now — do not retry-loop. Wait for the floating window to release calls.
409 ResourceExistsException Reused IngestionId. Retry with a timestamped ID.
404 ResourceNotFoundException Almost always the wrong Region. Then check the dataset ID.
429 ThrottlingException Backoff and retry.
401 AccessDeniedException Credentials / IAM.
500 InternalFailureException Retry once, then support case with the RequestId.

Check remaining budget:

aws quicksight list-ingestions --aws-account-id $ACCOUNT --data-set-id $DATASET \
  --region $REGION --max-results 100 \
  --query 'Ingestions[].RequestSource' --output text | tr '\t' '\n' | sort | uniq -c

Symptom: source database is being hammered by Quick Sight

Direct query only. The 2-minute visual timeout does not stop the query on Redshift — it keeps running and holds a WLM slot.

  1. Cancel the running queries on the cluster.
  2. Put Quick Sight in a dedicated WLM queue with a statement timeout.
  3. Longer term: move the aggregate workload to SPICE and keep direct query for drill-through only.

Escalating to AWS Support

Collect first:

aws quicksight list-ingestions  --aws-account-id $ACCOUNT --data-set-id $DATASET --region $REGION --max-results 100 > ingestions.json
aws quicksight describe-data-set --aws-account-id $ACCOUNT --data-set-id $DATASET --region $REGION > dataset.json
aws service-quotas list-service-quotas --service-code quicksight --region $REGION > quotas.json

Case body, in order: account ID · Region · dataset ID · ingestion ARN · ErrorInfo.Type verbatim · RequestId · UTC timestamps · RowInfo · what you ruled out and the commands you ran · one question with a factual answer.

Case type: Service limit increase only for quotas marked adjustable. Ingestion failures and INTERNAL_SERVICE_ERROR are Technical. Capacity charges are Account and billing.

Severity: stale dashboard = normal (12 h) or high (4 h) if genuinely business-critical. Not urgent. You can raise it later on Business Support+, Enterprise Support, or Unified Operations — but after raising to urgent or critical there's a 60-minute wait before the next change.

If the first reply is a doc link you've already read: name the page and the line, and state what remains unanswered. Reply on the existing case; never open a duplicate.


Escalation contacts

Fill this in for your organisation and keep it here.

Area Owner Channel
SPICE capacity / billing
Source schema changes
Networking / VPC connection
Secrets & credential rotation
AWS TAM (Enterprise Support)