Data/ML/AIDatabasesAround since 2014
Apache Spark
Also written as Spark, PySpark
A framework for processing very large datasets across many machines at once — the workhorse of big-data pipelines.
Think of it like
Like splitting a warehouse stocktake across fifty people with walkie-talkies instead of one person with a clipboard.
Junior or senior?
Junior sounds like
Has run PySpark jobs someone else wrote.
Senior sounds like
Talks about partitioning, shuffles or skew — the things that make a distributed job slow for non-obvious reasons.
Ask them
“Tell me about a Spark job that was slow. What turned out to be causing it?”