Data/ML/AIHigh signal

Distributed Training

Also written as Multi-GPU Training, Data Parallelism, Model Parallelism

Splitting one model's training across many machines because it will not fit, or will not finish, on a single one. Splitting the data is the common case; splitting the model itself is far harder.

Think of it like

Ten people reading one book a chapter each, then having to agree on the plot.

Junior or senior?

Genuinely differentiating, because most practitioners have never needed it.

Senior sounds like

Can describe a failure specific to the distributed setup rather than to the model.

Ask them

“What broke that would not have broken on a single machine?”