AWS Training
Modules Listen Certification
0:00 0:00

← SPICE Internals and Data at Scale

Starts this lesson and continues through 17 more to the end of the course.

Refresh strategy at scale

The decision you're actually making

Full refresh reloads everything. Incremental refresh reloads a window. That framing is correct and almost useless, because it hides the thing that actually determines whether your data is right:

Incremental refresh does not update anything outside the look-back window. Ever.

Not "slowly". Not "on the next run". Never — until something triggers a full refresh. If your data gets corrected retroactively, or your chosen date column isn't the column that changes when a row changes, incremental refresh will happily serve stale rows forever while reporting success.

That is the lesson. Everything else is scheduling mechanics.

What incremental refresh actually does

From Refreshing SPICE data (verified 2026-08-09):

"An incremental refresh queries only data defined by the dataset within a specified look-back window. It transfers all insertions, deletions, and modifications to the dataset, within that window's timeframe, from its source to the dataset. The data currently in SPICE that's within that window is deleted and replaced with the updates."

Read the mechanism carefully, because it is delete-and-replace, not merge:

  1. Query the source for rows where the chosen date column falls in the window.
  2. Delete the rows currently in SPICE whose date column falls in that window.
  3. Insert what was queried.

AWS's own worked example:

"let's say you have a dataset with 180,000 records that contains data from January 1 to June 30. On July 1, you run an incremental refresh on the data with a look-back window of seven days. Quick Sight queries the database asking for all data since June 24 (7 days ago), which is 7,000 records. Quick Sight then deletes the data currently in SPICE from June 24 and after, and appends the newly queried data. … Rather than having to ingest 180,000 records every day, it only has to ingest 7,000 records."

Availability: Enterprise edition only, and "for SQL-based data sources, such as Amazon Redshift, Amazon Athena, PostgreSQL, or Snowflake". Note "such as" — that list is illustrative, not exhaustive. I could not verify a complete supported-source list from the pages fetched for this lesson; check the console's Refresh type control for the dataset in question, which is the authoritative answer for your setup.

The three ways this goes wrong

⚠️ The following are my inferences from the documented delete-and-replace mechanism above, not statements AWS makes explicitly. They are, however, direct consequences of it, and each one is a real production failure I would design against.

1. You picked the wrong date column. If you window on created_at but rows are updated later, an update to a January row will never be seen — January is outside every future window. Window on the column that changes when the row changes (updated_at, modified_ts, a CDC watermark), not on the column that describes the event.

2. Late-arriving data outside the window is invisible. A seven-day window and a source system that back-posts corrections at 30 days means those corrections never land. The dashboard will be confidently wrong, and every ingestion will report COMPLETED.

3. Hard deletes outside the window never propagate. The mechanism deletes SPICE rows within the window. A row deleted at source with an old date stays in SPICE indefinitely.

The standard mitigation — and this is a recommendation, not documentation — is a periodic full refresh alongside the incremental schedule: incremental for freshness, a weekly or monthly full refresh for correctness. Budget it against the 32-call quota from lesson 2, and run it when nobody's looking at the dashboard.

Scheduling: the rules that constrain you

Verbatim constraints from the refresh page:

Rule Detail
Schedules per dataset (console) 5. "When you have created five, the Create button is turned off."
Timing accuracy "Scheduled dataset ingestions take place within 10 minutes of the scheduled date and time."
Full-refresh frequencies Daily, Weekly, Monthly (Standard + Enterprise); Hourly (Enterprise only)
Incremental frequencies Every 15 minutes, Every 30 minutes, Hourly, Daily, Weekly, Monthly (Enterprise)
Monthly on the 29th–31st Choose Last day of month
Hourly exclusivity (full) "If you decide to use an hourly refresh, you can't also use additional refresh schedules."
Sub-hourly exclusivity (incremental) Same rule for 15-minute, 30-minute, and hourly incremental schedules

⚠️ The 10-minute window is not a service-level guarantee you can build on. If a downstream job consumes a dashboard export at 06:00 and the refresh is scheduled for 05:55, you have a five-minute margin against a documented ten-minute variance. Schedule the refresh, then poll list-ingestions for COMPLETED — don't sleep and hope.

⚠️ Exclusivity is an ordering problem. To move from daily to hourly you must delete the daily schedule first; to move back you must delete the hourly first. There's no in-place switch, so a badly ordered automation leaves the dataset with no schedule at all. Automate it as delete-then-create, and verify.

Deleting an incremental configuration triggers a full refresh

"Deleting an incremental refresh configuration starts a full refresh. As part of this full refresh, all the configurations prepared for incremental refreshes are removed."

On a two-billion-row dataset that is not a config change — it is a multi-hour reload, one of your 32 daily ingestion calls, and a load spike on the source. Do it deliberately, in a window, not on a Friday afternoon.

Schema changes break refreshes, and nothing auto-heals

Repeating the sentence from lesson 3 because it belongs to refresh strategy as much as to troubleshooting:

"If there is a schema change in a database, Quick Sight will not be able to auto-detect it, resulting in an ingestion failure. Edit and save the dataset to update the schema and avoid ingestion failures."

There is no API-free way around this. Put "re-save affected Quick Sight datasets" in your data platform's schema-change runbook, or accept that every upstream ALTER TABLE is a scheduled outage of your dashboards.

S3 datasets and the manifest

For S3-backed datasets, a refresh re-reads a manifest, and you choose which:

The useful pattern: keep the manifest at a stable S3 URL and rewrite its contents from your pipeline. Quick Sight keeps pointing at the same URL, picks up the new file list on every refresh, and nobody touches the console. Remember the ceiling — 1,000 files per manifest (lesson 2).

Choosing a strategy

Situation Strategy
Small dataset (< a few GB), source is cheap to query Full refresh, daily or hourly. Simplest thing that works.
Large, append-mostly, reliable updated_at Incremental on updated_at + weekly full refresh
Large, with retroactive corrections beyond the window Widen the window to cover the correction horizon, or accept a nightly full refresh
Needs sub-15-minute freshness Not SPICE. See lesson 5.
Source is per-query priced (Athena) Strongly prefer SPICE; each direct-query dashboard view is a billed scan

The opinionated default: start with a full refresh. Move to incremental only when a full refresh stops fitting its window or costs too much — and when you move, pick the date column deliberately, write down your correction horizon, and schedule the periodic full refresh in the same change. An incremental schedule without a full-refresh companion is a correctness bug waiting for a quarter-end.

Check yourself

  1. Incremental refresh, 7-day window on order_date. Finance corrects a March order in August. When does the dashboard show the correction?
  2. You need to move a dataset from a daily to an hourly schedule via automation. What's the order of operations and what breaks if you get it wrong?
  3. Why is "the refresh is scheduled for 05:55, the export runs at 06:00" fragile?
  4. A dataset has 4 schedules and you want a 15-minute incremental. What must happen first?
  5. You delete an incremental refresh configuration on a 1.8-billion-row dataset at 16:45 on a Friday. What have you just started?
Answers
  1. Never — until a full refresh runs. March is outside every future 7-day window, so the corrected row is never re-queried. This is the core hazard of incremental refresh.
  2. Delete the daily schedule, then create the hourly one. Hourly is exclusive of all other schedules. Create-then-delete fails, and a crash between the two steps leaves the dataset with no schedule — silently stale.
  3. Ingestions run "within 10 minutes of the scheduled date and time". A 5-minute margin against a documented 10-minute variance loses regularly. Poll list-ingestions for COMPLETED instead.
  4. Nothing about the count of 5 — you're under it. But sub-hourly incremental schedules are exclusive, so all four existing schedules must be deleted first.
  5. A full refresh of 1.8 billion rows, consuming an ingestion call, hammering the source, and running into the weekend — plus the removal of every incremental configuration on that dataset.

Teaching this section

← PreviousIngestion mechanics and the failure taxonomyNext →When SPICE is the wrong answer — and talking to AWS Support