Data/ML/AIHigh signal

Pretraining

Also written as Pre-training, Base Model

The first, enormously expensive stage of building a large model: training it on a huge slice of the internet, books and code so it learns language and general knowledge. Fine-tuning and RLHF come afterwards to make it helpful and safe. The raw result is called a base model.

Think of it like

Years of general schooling before anyone trains you for a specific job.

Junior or senior?

Very few people have genuinely pretrained a large model, since it happens at a handful of AI labs and big tech companies. Most candidates who mention it fine-tuned a pretrained model, which is a far smaller job.

Senior sounds like

Can talk about data mixture, cluster scale and training runs that failed partway through.

Ask them

“Were you training a model from scratch or adapting one someone else had pretrained? Roughly how many GPUs and how long did a run take?”

Sounds like real experience

Gives rough but specific scale (hundreds or thousands of GPUs, weeks of training) and describes a real problem such as loss spikes, bad data or hardware failures mid-run.

Probe further if

Describes downloading a model from Hugging Face and fine-tuning it, while calling that pretraining.