Security and governance

Anthropomorphism

stable definition

Definition

Anthropomorphism is attributing human mental states, motives, understanding, emotions, or agency to an AI model or system because its language or behavior appears human-like. Conversational fluency makes this especially easy with generative systems.

Why it matters

Human metaphors can make interfaces understandable, but they can also distort responsibility and risk judgments. Saying a model "knows," "wants," "decides," or "refuses" may hide the roles of training, prompts, tools, policies, operators, and stochastic inference. Users may overtrust confident language, disclose more information, or assume stable intentions that the system does not possess.

Distinguish it from nearby terms

Agency is an operational property of a system's ability to pursue goals and take actions. Anthropomorphism is an interpretation imposed by people. A system can have consequential agency without having human-like experience or intent.

Check your understanding

Replace the human-state claim with an observable mechanism: what input, model behavior, system rule, tool, or operator action produced the outcome?