Definition
AI safety is the field and practice of keeping AI-system risks within acceptable bounds across design, deployment, operation, and retirement. It asks which hazards a system can create, who may be harmed, how severe and likely the harm is, which controls reduce it, how failure will be detected, and how the system can recover or stop.
The field covers immediate engineering failures and broader risks: unreliable outputs, unsafe tool actions, misuse, loss of control, human overreliance, discrimination, privacy harm, security compromise, systemic dependence, and failures from capabilities or environments that change after launch. Different communities emphasize different parts of that range, so safe AI is not a complete claim without a domain and risk threshold.
Safety is a continuing case
Pre-release evaluation cannot cover every real condition. Safety work links a bounded assurance claim to evidence, operating limits, monitoring, incident response, change management, and conditions that invalidate the claim. The responsible organization still owns that case when a model or platform comes from a vendor.
Distinguish it from nearby terms
AI security focuses on adversaries, unauthorized access, and compromise. Alignment asks whether behavior remains compatible with intended goals and values. Responsible AI includes safety along with governance, fairness, transparency, privacy, and accountability. The boundaries overlap, but none substitutes for a concrete hazard analysis.
Check your understanding
A medical summarizer is accurate on average but occasionally drops allergy information. Is it safe? Average accuracy is insufficient. Define the hazardous scenario, affected patient, severity, operating conditions, detection and escalation control, residual risk, and who may authorize use.