Data/ML/AIHigh signalAround since 2017

RLHF

Also written as Reinforcement Learning from Human Feedback

A training technique where human reviewers rank AI outputs, and the model is further trained to produce more of what humans preferred — a key step in making raw models helpful and safe.

Think of it like

Like a chef adjusting recipes based on customer ratings over and over, gradually cooking more of what diners actually rate highly.

Junior or senior?

Junior sounds like

Has only consumed already-trained models.

Senior sounds like

Has worked on the human-feedback or labeling side of a training pipeline.

Ask them

“Have you worked on the human-feedback or labeling side of a training pipeline, or mainly consumed already-trained models?”