"AI Safety": Alignment vs. Guardrails vs. Red Teaming
New'AI safety experience' covers three different kinds of work. Alignment is research into making models themselves behave as intended, usually at AI labs and often PhD-level. Guardrails are engineering added around a product to block harmful, off-topic or leaked output. Red teaming is deliberately attacking an AI system — jailbreaks, prompt injection — to find weaknesses before others do, and sits close to security work.
How to tell them apart on a resume
Alignment (research)
RLHF, constitutional AI, interpretability, scalable oversight, published papers — at an AI lab or research institute.
Guardrails (product engineering)
Content filters, input and output checks, PII redaction, NeMo Guardrails, Llama Guard, moderation APIs — shipping AI features safely.
Red teaming (security)
Jailbreaks, prompt injection, adversarial testing, garak, PyRIT, security background, bug bounty — breaking AI systems on purpose.
The question that settles it
“Was your safety work changing how the model itself behaves, adding checks around a product, or attacking AI systems to find weaknesses?”
Read the full definitions
Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.