AWS Training
Modules Listen Certification
0:00 0:00

← PySpark

PS0 Quiz — PySpark

18 questions. Four options each, one correct answer, no partial credit.


1. After df = df.filter(col("x") > 5) runs, what has happened?

2. Your traceback names the write line, but the bug is a bad input path. Why?

3. Which of these does NOT move data into the driver?

4. Apache Arrow in PySpark is:

5. The documented default of spark.sql.execution.arrow.maxRecordsPerBatch is:

6. The biggest hidden cost of a Python UDF inside a filter() is:

7. Which pandas UDF type loads all data for a group into memory?

8. Which pandas Function API can return an arbitrary number of output rows?

9. A DataFrame is consumed by count(), then collect(), then write. Without caching:

10. 199 tasks finish fast, one runs 40 minutes. Doubling the cluster will:

11. A single 5 GB GZIP input file causes:

12. coalesce(1) after an expensive transformation:

13. Salting a hot join key costs you:

14. Storage partitions (partitionBy) versus Spark partitions:

15. In tests, setting spark.sql.shuffle.partitions to 2:

16. write.mode("append") in a job that an orchestrator may retry:

17. A print() inside a UDF appears in:

18. Before trusting any default quoted in this module on Glue or EMR, you should:


Answers and explanations

1 — B. Transformations build a logical plan; nothing executes until an action. This is why a script can appear to run instantly and then spend all its time on the last line.

2 — B. Everything before an action is plan construction and cannot fail on unread data. The action triggers execution, so the error surfaces there. Debug backwards from the action.

3 — C. write.parquet() is distributed — executors write directly to storage. The other three all bring rows to the driver, though take(10) is bounded and therefore usually safe.

4 — B. Verbatim: "an in-memory columnar data format that is used in Spark to efficiently transfer data between JVM and Python processes." It narrows the boundary cost; it doesn't remove the boundary.

5 — B. 10,000 records per batch, per the current documentation.

6 — C. The UDF is opaque to Catalyst, so the predicate can't be pushed to the scan. Spark reads everything and filters in Python, and that extra I/O usually dwarfs the UDF's own cost. A is a real but minor factor; B is wrong — the Python process sits outside the JVM heap.

7 — C. Series to Scalar: documented as not supporting partial aggregation, with "all data for a group or window loaded into memory", and only unbounded windows supported. That turns skew into an OOM.

8 — B. mapInPandas is documented as able to "return the output of arbitrary length". Both applyInPandas variants load a whole group (or cogroup) into memory and don't have that property.

9 — B. A DataFrame is a plan, not a result. Each action re-executes from source. Cache when more than one action consumes it — and unpersist afterwards.

10 — B. This is skew. Extra executors have nothing to do; one partition holds the work. Fix the distribution instead of the capacity.

11 — B. GZIP is not splittable, so the file must be read by a single task regardless of configuration. A common cause of apparently single-threaded jobs.

12 — B. coalesce avoids the shuffle by narrowing upstream parallelism, so the expensive computation can run in one task. repartition(1) pays for a shuffle but preserves upstream parallelism — the correct choice when you need one file and parallel compute.

13 — B. The replicated side is multiplied by the salt factor N, which is real extra shuffle and compute. Salt only genuinely hot keys, and only when filtering and broadcasting aren't available.

14 — B. Spark partitions are in-memory parallelism units (one task each); storage partitions are directories created by partitionBy. They interact through file pruning but are distinct concepts.

15 — B. The default creates a large number of mostly-empty partitions per shuffle, which dominates runtime for five-row test data. It doesn't reduce correctness coverage.

16 — B. Orchestrators, spot reclamation, Glue and Step Functions all retry. Append-mode writes the data again. Assume every job runs at least twice and design for idempotency.

17 — B. UDFs execute on executors, so their output goes to executor logs — a different location from driver logs. This is the usual reason people conclude their UDF "isn't running".

18 — B. Glue and EMR pin Spark versions behind the current release, and these defaults have moved between versions. Print spark.version in the actual environment and read the matching docs.

Scoring

Score Reading
16–18 Strong. Rehearse the interview file out loud, then move to GL0 (Glue).
12–15 Good. Re-read lessons 3 and 4 — UDF cost and skew carry most of the value.
8–11 Do the lab. Parts 2, 3 and 4 in particular.
≤ 7 Start at lesson 1. Everything here is a consequence of the boundary.

The five that matter most in production: 6 (pushdown loss), 9 (plan not result), 10 (skew ≠ capacity), 12 (coalesce narrowing), 16 (idempotency). Each one costs real money or real data.