DPO / Preference Tuning
Also written as DPO, Direct Preference Optimization, Preference Optimization, Preference Tuning
Teaching a model to give the kind of answers people prefer by showing it pairs of responses labelled 'better' and 'worse'. DPO is a simpler, cheaper alternative to RLHF that has become a common way to do this step.
Think of it like
Training a new writer by showing them two drafts side by side and saying which one the editor liked, over and over.
Junior or senior?
Hands-on experience here usually means real post-training work, beyond basic fine-tuning.
Senior sounds like
Can talk about where the preference data came from and how they checked the model actually improved rather than just changing tone.
Ask them
“Where did your preference pairs come from, and how did you tell whether the tuned model was really better?”