AWS Training
Modules Listen Certification

← SPICE Internals and Data at Scale

Starts this lesson and continues through 16 more to the end of the course.

When SPICE is the wrong answer — and talking to AWS Support

Part 1 — Admitting SPICE is wrong

There is a point where the correct engineering answer stops being "tune the ingestion" and becomes "stop importing this data".

The signals, in rough order of how definitive they are:

Signal What it means
A single dataset needs > 2 billion rows or > 2 TB Hard stop. Not adjustable. Not negotiable.
You need freshness under ~15 minutes You're fighting the 32-call quota and the 10-minute schedule variance
Full refresh no longer fits its window, and incremental can't be made correct You're choosing between stale and late
The source is already a fast, columnar, well-tuned warehouse You're caching a cache
Users need row-level detail on billions of rows SPICE is an aggregation engine; you're using it as a table viewer

The four honest options when you exceed a dataset quota

1. Aggregate before ingesting. The most under-used option, and usually the right one. A dashboard showing daily revenue by region does not need 2 billion order-line rows; it needs a pre-aggregated table. Push the aggregation into Redshift, Athena, or your transformation layer. A 2-billion-row fact table often becomes a 5-million-row summary that fits in SPICE with room to spare and refreshes in minutes.

2. Split by dimension, not by size. Separate datasets per region, per business unit, per year. This is legitimate — but remember lesson 2: splitting multiplies your CreateIngestion consumption, because the 32-call quota is per Region, not per dataset. Four datasets on an hourly schedule is 96 calls against a quota of 32. Split by dimension only if you can also drop the cadence.

3. Direct query for the deep detail, SPICE for the summary. The pattern that works in practice: the landing dashboard runs on a small SPICE summary and is instant; a drill-through action opens a detail sheet backed by direct query, which only a handful of users hit at a time. You get SPICE economics on the 95% of traffic that's aggregate, and full fidelity where it's actually needed.

4. Move the workload out of Quick Sight. If the requirement is "export 400 million rows", that isn't a BI requirement. It's an UNLOAD to S3.

Direct query: what you take on

From lesson 2, restated because it's the deciding factor:

Design rule for direct query at scale: a dedicated WLM queue with a statement timeout on Redshift, result caching where the engine offers it, and a hard cap on how many concurrent viewers a detail sheet can have. Direct query moves the load onto your warehouse — budget for it there.

Part 2 — Talking to AWS Support

Before you open anything: which of three things is this?

AWS Support's console offers three case types (Case management, verified 2026-08-09):

Case type Available to Use it for
Account and billing All customers SPICE capacity charges, subscription/pricing questions
Service limit increase All customers Only quotas marked adjustable in the Service Quotas table
Technical Not Basic Support Ingestion failures, INTERNAL_SERVICE_ERROR, behaviour questions

⚠️ Do not file a service limit increase for a non-adjustable quota. The 2-billion-row per-dataset quota and the 32-call CreateIngestion quota are both marked not adjustable. Filing against them burns days waiting for a "no". Check the Adjustable column first — that's the entire purpose of the CLI command in lesson 2.

Severity: pick honestly

The documented severity levels and first-response times:

Severity Code First response Applies when
General guidance low 24 hours A development question or feature request
System impaired normal 12 hours Non-critical functions abnormal, or a time-sensitive dev question
Production system impaired high 4 hours Important functions impaired or degraded
Production system down urgent 1 hour Business significantly impacted, important functions unavailable
Business-critical system down critical < 30 min (Business Support+) · < 15 min (Enterprise Support) · 5 min from an Incident Management Engineer (Unified Operations) Business at risk, critical functions unavailable

AWS's own guidance:

"it's a best practice to choose the highest severity only for cases that can't be worked around or that directly affect production applications."

A stale executive dashboard is normal, occasionally high. It is not urgent. Inflating severity costs you credibility on the case where you genuinely need it — and on Business Support+, Enterprise Support, or Unified Operations you can raise the severity later if it escalates, so there's no reason to over-call it at the start.

⚠️ Note the re-escalation waiting period: after raising severity, you may have to wait out the first-response window of the new level before raising it again, and raising to urgent or critical imposes a 60-minute wait before the next change.

The evidence pack

"Make your description as detailed as possible. Include relevant resource information … For example, to troubleshoot performance, include timestamps and logs."

For a Quick Sight SPICE case, "relevant resource information" is specific. Gather all of this before you open the case:

ACCOUNT=111122223333
DATASET=sales-fact-v3
REGION=us-east-1

# 1. The failing ingestion history — the single most useful artefact
aws quicksight list-ingestions \
  --aws-account-id $ACCOUNT --data-set-id $DATASET \
  --max-results 100 --region $REGION > ingestions.json

# 2. The dataset definition (import mode, columns, size expectations)
aws quicksight describe-data-set \
  --aws-account-id $ACCOUNT --data-set-id $DATASET \
  --region $REGION > dataset.json

# 3. Your applied quotas — proves you checked
aws service-quotas list-service-quotas \
  --service-code quicksight --region $REGION > quotas.json

Then the case body. Include, in this order:

  1. AWS account ID and Region — Region is wrong or missing in a startling number of cases, and Quick Sight is intensely Region-scoped.
  2. Dataset ID and Ingestion ARN of a specific failed run.
  3. ErrorInfo.Type — the enum value, verbatim. Not your paraphrase of the console message.
  4. RequestId from the failed API call, if you have one. This is the highest-leverage single field in the case.
  5. Timestamps in UTC, with the range you want them to look at.
  6. RowInfo — ingested, dropped, total.
  7. What you already ruled out, with the commands you ran.
  8. The specific question, phrased so it has a factual answer.

Ask answerable questions

This is where most cases go wrong. Compare:

❌ Unanswerable ✅ Answerable
"Our refreshes keep failing, please advise." "Ingestion arn:…:ingestion/abc returned ErrorInfo.Type = ACCOUNT_CAPACITY_LIMIT_EXCEEDED at 2026-08-09T03:14Z. Account SPICE usage reads 1.9 TB of 2 TB. Confirm the account-level capacity at that timestamp."
"Can you increase our SPICE limit?" "Confirm whether the 2-billion-row per-dataset quota is adjustable. Service Quotas shows no entry for it; the user guide documents it as a fixed quota."
"Is our data too big for QuickSight?" "For a 3-billion-row fact table, confirm the supported patterns: (a) pre-aggregation, (b) dataset partitioning by dimension, (c) direct query. Is there a documented guidance page for datasets above the 2-billion-row quota?"
"Refreshes are slow." "IngestionTimeInSeconds on dataset X has grown from 2,100 to 9,400 over 60 days with row count flat at 400 M. Is there a known regression, or is this expected as SPICE capacity utilisation approaches 90%?"

The three questions worth asking about large data

If your problem is "we have far more data than this thing wants", these are the questions that actually produce new information rather than a doc link:

  1. "Does a scheduled refresh consume the API_CREATE-INGESTION quota?" The docs describe the quota two different ways and never state this. Ask for it in writing; it's a one-sentence answer that changes refresh architecture.
  2. "What is the supported architecture above 2 billion rows per dataset?" Not "can you raise it" — you already know the answer. Ask what AWS's reference pattern is. Support has visibility into what large customers actually do.
  3. "For our account, in this Region, what is the current bundled vs purchased capacity split, and what does auto-purchase do at our usage curve?" This is a billing-adjacent question with a factual answer, and it's the one that determines whether your cost model survives the year.

Escalating well

Check yourself

  1. Your dataset needs 3 billion rows. Which case type do you open?
  2. Nightly refresh fails; the exec dashboard is stale by 09:00. What severity, and why not higher?
  3. What is the single highest-leverage field to include in a technical case?
  4. Give a reason not to solve a dataset-quota problem by splitting into four datasets.
  5. Direct query on Redshift: a visual times out at 2 minutes. What has not happened?
Answers
  1. None for a limit increase — the quota is not adjustable. If you want architectural guidance, open a Technical case (or use your TAM) asking for the supported pattern above 2 billion rows.
  2. normal (System impaired), possibly high if it's genuinely business-critical. Not urgent: AWS guidance reserves the top levels for things that can't be worked around or that take production down, and you can raise severity later if impact grows.
  3. The RequestId from the failed API call — closely followed by the exact ErrorInfo.Type enum value. Both let Support look up the actual event rather than reason from your description.
  4. Because the 32-call CreateIngestion quota is per Region, not per dataset. Four datasets on an hourly schedule is 96 calls against a quota of 32.
  5. The query has not stopped. The Redshift driver doesn't react to the client-side timeout, so it keeps running and holding a WLM slot until it completes.

Teaching this section

← PreviousRefresh strategy at scaleFinished →Cheat sheet, lab & quiz