Security and governance

AI red teaming

stable definition

Definition

AI red teaming is structured adversarial testing intended to discover failure, misuse, exploitation, or constraint violations before attackers or users do.

Distinguish it from nearby terms

Ordinary evaluation measures expected behavior; red teaming actively searches for unexpected harmful behavior and attack paths.

Check your understanding

State the threat model, attacker access, success criteria, and how findings become durable controls.