AWS Training
Modules Listen Certification

← SPICE Internals and Data at Scale

Q2 Cheat sheet — SPICE at scale

All figures verified 2026-08-09 against the AWS General Reference quota table and the Quick user guide. Re-verify before quoting in a support case.

Sizing formula

Logical size (bytes) =
  (numeric cells × 12)
+ (date cells    × 12)
+ Σ (24 + UTF-8 byte length) per text cell

Computed after type transforms and calculated columns, at dataset save. Changes made in an analysis do not change SPICE size. Lat/long are numeric; all other geospatial categories are strings.

⚠️ Single dataset only. Do not sum these to estimate account capacity — read the admin SPICE capacity page.

Rule of thumb: a text column costs ≥ 24 bytes/row before its content. Dropping one 20-byte text column from 1 B rows saves ≈ 44 GB.

Quotas — the adjustable column is the one that matters

Quota Value Adjustable
Rows / dataset — Standard 25,000,000 ❌
Size / dataset — Standard 25 GB ❌
Rows / dataset — Enterprise 2,000,000,000 ❌
Size / dataset — Enterprise 2 TB ❌
CreateIngestion calls / 24h — Enterprise 32 ❌
CreateIngestion calls / 24h — Standard 8 ❌
Columns per file / fields per dataset 2,000 ❌
Column name length 127 Unicode chars ❌
Field length (classic / new data prep) 2,047 / 65,534 Unicode chars ❌
Files per S3 manifest 1,000 ❌
Visual query timeout 120 s ❌
Dataset preview timeout 45 s ❌
Calculated field expression 250,000 chars ❌
Custom actions per visual 10 ❌
Email aliases per group (reports) 5,000 ❌
Schedules per dataset (console) 5 —

Rows OR size — whichever you hit first. Row-level security does not reduce footprint. CreateIngestion window is floating 24h, not calendar.

aws service-quotas list-service-quotas --service-code quicksight --region us-east-1 \
  --query 'Quotas[?Adjustable==`false`].{Name:QuotaName,Value:Value}' --output table

Service code is quicksight, not quick. Not in the table ⇒ not adjustable.

Two ceilings — never confuse them

Per-dataset quota (2 B rows / 2 TB) Account + Region capacity
❌ Not adjustable ✅ Buy more
→ redesign → purchase order

First-line diagnosis — run this before anything else

aws quicksight list-ingestions \
  --aws-account-id $ACCOUNT --data-set-id $DATASET --region $REGION --max-results 100 \
  --query 'Ingestions[].{When:CreatedTime,Status:IngestionStatus,Src:RequestSource,
      Type:RequestType,Secs:IngestionTimeInSeconds,In:RowInfo.RowsIngested,
      Dropped:RowInfo.RowsDropped,Err:ErrorInfo.Type}' --output table

IngestionStatus: INITIALIZED | QUEUED | RUNNING | FAILED | COMPLETED | CANCELLED RequestSource: MANUAL | SCHEDULED · RequestType: INITIAL_INGESTION | EDIT | INCREMENTAL_REFRESH | FULL_REFRESH

⚠️ Alert on RowsDropped > 0, not just FAILED. COMPLETED + dropped rows = silently wrong dashboard. ⚠️ RequestType: EDIT ⇒ someone saved the dataset and triggered a reload.

Starting an ingestion

aws quicksight create-ingestion --aws-account-id $ACCOUNT --data-set-id $DATASET \
  --ingestion-id "manual-$(date +%Y%m%dT%H%M%S)" \
  --ingestion-type FULL_REFRESH --region $REGION

IngestionType: FULL_REFRESH | INCREMENTAL_REFRESH. IngestionId 1–128 chars ^[a-zA-Z0-9-_]+$. Always timestamp it — it's a PUT and ResourceExistsException is real.

HTTP Exception Retry?
401 AccessDeniedException No
400 InvalidParameterValueException No
409 LimitExceededException No — terminal for the day
409 ResourceExistsException New ID
404 ResourceNotFoundException Check Region
429 ThrottlingException Yes, backoff
500 InternalFailureException Once, then case

Error triage — the 6 to know cold

ErrorInfo.Type Action
ACCOUNT_CAPACITY_LIMIT_EXCEEDED Buy / release capacity
DATA_SET_SIZE_LIMIT_EXCEEDED Redesign — not adjustable
SQL_SCHEMA_MISMATCH_ERROR Edit + save the dataset to refresh its schema
UNRESOLVABLE_HOST / UNROUTABLE_HOST DNS / routing + SGs — different fixes
INGESTION_SUPERSEDED, REFRESH_SUPPRESSED_BY_EDIT Not a defect — don't page
INTERNAL_SERVICE_ERROR Retry once → support case with RequestId

Full 45-value enum: ErrorInfo

Refresh

Incremental = delete-and-replace within the window. Query window → delete SPICE rows in window → insert. Enterprise only, SQL-based sources.

⚠️ Nothing outside the window is ever updated. Late corrections and hard deletes with old dates are invisible until a full refresh. Pair every incremental schedule with a periodic full refresh.

Rule Value
Schedules per dataset 5
Timing accuracy within 10 minutes of scheduled time
Full refresh Daily / Weekly / Monthly · Hourly = Enterprise only
Incremental 15 min / 30 min / Hourly / Daily / Weekly / Monthly
Hourly & sub-hourly Exclusive — delete other schedules first
Delete incremental config Triggers a full refresh and wipes incremental config
29th–31st monthly Use "Last day of month"

Window on the column that changes when the row changes (updated_at), not on the event date.

Capacity

Per Region, shared by all users in the account. Bundled (not releasable) + Purchased (releasable if unused). Release order: remove/convert datasets first, then release.

Auto-purchase: console-created accounts on; API-created and pre-existing off. No auto-decrement. Off + over capacity ⇒ refreshes fail. Turning off converts current usage to purchased capacity.

⚠️ The Quick console Region selector is independent of the AWS console's. Read the Region from the URL before buying.

Support case checklist

Account ID · Region · Dataset ID · Ingestion ARN · ErrorInfo.Type verbatim · RequestId · UTC timestamps · RowInfo · what you ruled out · one answerable question.

Case type For
Service limit increase Adjustable quotas only
Technical Ingestion failures, INTERNAL_SERVICE_ERROR, behaviour
Account and billing Capacity charges, subscriptions
Severity Code First response
General guidance low 24 h
System impaired normal 12 h
Production system impaired high 4 h
Production system down urgent 1 h
Business-critical system down critical <30 m Business+ · <15 m Enterprise · 5 m Unified Ops

Stale dashboard = normal/high. Not urgent. You can raise it later on paid plans.

Direct query gotcha

"not all database drivers react to the 2-minute timeout, for example Amazon Redshift … the query runs for as long as it takes for the response to return"

Dedicated WLM queue + statement timeout. The client-side timeout does not protect the cluster.