Data/ML/AISecurityEmergingAround since 2023
AI Red-Teaming
Also written as Red-Teaming (AI), Jailbreak Testing, Adversarial Testing (AI)
Deliberately trying to make an AI model misbehave — produce harmful content, leak data, ignore its instructions — before real users find the same holes.
Think of it like
Hiring someone to try to break into a building at night so you can fix the locks before an actual burglar tries.
Junior or senior?
Distinct from traditional security red-teaming, which targets infrastructure rather than model behavior.
Senior sounds like
Describes a specific technique that got a model to misbehave, and what the actual fix was.
Ask them
“What's a specific way you got a model to do something it shouldn't have, and how was it actually fixed?”