Definition
An input strategy intended to make a model bypass or disregard its trained or instructed safety restrictions.
Distinguish it from nearby terms
Jailbreaking targets safety policy; prompt injection more broadly redirects system behavior and may exploit tools or data.
Check your understanding
A model-level refusal does not replace application-level authorization and containment.