Data/ML/AIDatabasesAround since 2014

Apache Spark

Also written as Spark, PySpark

A framework for processing very large datasets across many machines at once — the workhorse of big-data pipelines.

Think of it like

Like splitting a warehouse stocktake across fifty people with walkie-talkies instead of one person with a clipboard.

Junior or senior?

Junior sounds like

Has run PySpark jobs someone else wrote.

Senior sounds like

Talks about partitioning, shuffles or skew — the things that make a distributed job slow for non-obvious reasons.

Ask them

“Tell me about a Spark job that was slow. What turned out to be causing it?”