Reinforcement Learning: Classic RL vs. RLHF vs. RLVR

New

How to tell them apart on a resume

Classic RL

Gym, PPO, DQN, simulations, robotics, game AI, recommendation or bidding systems — reward comes from an environment.

RLHF / preference tuning

Human preference data, reward models, DPO, PPO on LLMs, annotation pipelines — making a model more helpful and safe.

RLVR (verifiable rewards)

Reasoning models, maths or code tasks, automatic graders, GRPO, post-training at an AI lab — reward comes from a checkable answer.

The question that settles it

“In your reinforcement learning work, where did the reward come from — an environment or simulator, human preferences, or an automatic check of the answer?”

Read the full definitions

Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.