Reinforcement Learning: Classic RL vs. RLHF vs. RLVR
New'Reinforcement learning' on a resume can mean two separate worlds. Classic RL trains an agent by trial and error in an environment — games, robots, trading or ad bidding. In LLMs, RL is used after pre-training: RLHF rewards answers people preferred, and RLVR (the method behind reasoning models) rewards answers that can be automatically checked, like correct maths or passing code. The maths is related; the day-to-day work and employers are quite different.
How to tell them apart on a resume
Classic RL
Gym, PPO, DQN, simulations, robotics, game AI, recommendation or bidding systems — reward comes from an environment.
RLHF / preference tuning
Human preference data, reward models, DPO, PPO on LLMs, annotation pipelines — making a model more helpful and safe.
RLVR (verifiable rewards)
Reasoning models, maths or code tasks, automatic graders, GRPO, post-training at an AI lab — reward comes from a checkable answer.
The question that settles it
“In your reinforcement learning work, where did the reward come from — an environment or simulator, human preferences, or an automatic check of the answer?”
Read the full definitions
Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.