Hadoop vs. Spark vs. Databricks

How to tell them apart on a resume

Hadoop

HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Oozie, Cloudera, Hortonworks — on-premises clusters, usually 2010s work.

Spark

PySpark, Spark SQL, DataFrames, Spark Streaming, EMR, Dataproc, performance tuning — large-scale processing code the engineer wrote themselves.

Databricks

Delta Lake, Unity Catalog, notebooks, workflows, lakehouse, Databricks certifications — the managed cloud version, usually current work.

The question that settles it

“Was your big-data work on an on-premises Hadoop cluster, Spark code you wrote and tuned, or the Databricks platform — and what was the largest dataset you processed?”

Read the full definitions

Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.