"AI Safety": Alignment vs. Guardrails vs. Red Teaming

New

How to tell them apart on a resume

Alignment (research)

RLHF, constitutional AI, interpretability, scalable oversight, published papers — at an AI lab or research institute.

Guardrails (product engineering)

Content filters, input and output checks, PII redaction, NeMo Guardrails, Llama Guard, moderation APIs — shipping AI features safely.

Red teaming (security)

Jailbreaks, prompt injection, adversarial testing, garak, PyRIT, security background, bug bounty — breaking AI systems on purpose.

The question that settles it

“Was your safety work changing how the model itself behaves, adding checks around a product, or attacking AI systems to find weaknesses?”

Read the full definitions

Open the full tool for the other look-alike pairs, role profiles, and the JD decoder.