Data/ML/AISecurityEmergingAround since 2023

AI Red-Teaming

Also written as Red-Teaming (AI), Jailbreak Testing, Adversarial Testing (AI)

Deliberately trying to make an AI model misbehave — produce harmful content, leak data, ignore its instructions — before real users find the same holes.

Think of it like

Hiring someone to try to break into a building at night so you can fix the locks before an actual burglar tries.

Junior or senior?

Distinct from traditional security red-teaming, which targets infrastructure rather than model behavior.

Senior sounds like

Describes a specific technique that got a model to misbehave, and what the actual fix was.

Ask them

“What's a specific way you got a model to do something it shouldn't have, and how was it actually fixed?”