Data/ML/AIHigh signalAround since 2017
RLHF
Also written as Reinforcement Learning from Human Feedback
A training technique where human reviewers rank AI outputs, and the model is further trained to produce more of what humans preferred — a key step in making raw models helpful and safe.
Think of it like
Like a chef adjusting recipes based on customer ratings over and over, gradually cooking more of what diners actually rate highly.
Junior or senior?
Junior sounds like
Has only consumed already-trained models.
Senior sounds like
Has worked on the human-feedback or labeling side of a training pipeline.
Ask them
“Have you worked on the human-feedback or labeling side of a training pipeline, or mainly consumed already-trained models?”