Hadoop vs. Spark vs. Databricks
Three big-data names that look like one skill but mark different generations. Hadoop (2006) stores data across many machines and processes it slowly from disk; it is now mostly legacy. Spark (2014) processes data in memory and is much faster — it can run on Hadoop or without it. Databricks is the company founded by Spark's creators, selling a managed cloud platform that runs Spark. A 'Hadoop' resume needs a modern update; a 'Databricks' resume can still be shallow if it is only notebooks.
How to tell them apart on a resume
Hadoop
HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Oozie, Cloudera, Hortonworks — on-premises clusters, usually 2010s work.
Spark
PySpark, Spark SQL, DataFrames, Spark Streaming, EMR, Dataproc, performance tuning — large-scale processing code the engineer wrote themselves.
Databricks
Delta Lake, Unity Catalog, notebooks, workflows, lakehouse, Databricks certifications — the managed cloud version, usually current work.
The question that settles it
“Was your big-data work on an on-premises Hadoop cluster, Spark code you wrote and tuned, or the Databricks platform — and what was the largest dataset you processed?”
Read the full definitions
Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.