Data/ML/AI

DPO / Preference Tuning

Also written as DPO, Direct Preference Optimization, Preference Optimization, Preference Tuning

Teaching a model to give the kind of answers people prefer by showing it pairs of responses labelled 'better' and 'worse'. DPO is a simpler, cheaper alternative to RLHF that has become a common way to do this step.

Think of it like

Training a new writer by showing them two drafts side by side and saying which one the editor liked, over and over.

Junior or senior?

Hands-on experience here usually means real post-training work, beyond basic fine-tuning.

Senior sounds like

Can talk about where the preference data came from and how they checked the model actually improved rather than just changing tone.

Ask them

“Where did your preference pairs come from, and how did you tell whether the tuned model was really better?”