Definition
Claude Sonnet is Anthropic's model line positioned to balance capability, response speed, and cost for everyday coding and knowledge work. Sonnet is a family of releases; a full name such as Claude Sonnet 5.5 identifies one model in that line.
The position is relative to the rest of Anthropic's range. It does not define a universal size, architecture, or minimum reasoning ability.
Origin and attribution
Anthropic announced Sonnet with Opus and Haiku in the Claude 3 family on March 4, 2024. Sonnet was available on the launch date. The announcement credits Anthropic and does not identify a sole creator of the line.
Claude Sonnet 5.5 launched on September 28, 2026. Anthropic describes it as a faster, lower-cost complement to Opus 5.5 for well-scoped everyday tasks, including fixing bugs and creating documents. The Claude API identifier is claude-sonnet-5-5.
Scope and limits
Treat statements about the line's balance of speed and intelligence as product positioning that requires task evaluation. Anthropic's Sonnet 5.5 announcement notes that similar benchmark scores can coexist with differences on complex, open-ended work requiring sustained judgment. One score cannot establish equivalence across the whole application.
Supported tools, thinking settings, and safeguards belong to individual releases. Sonnet 5.5 introduces cyber safeguards and fallbacks that older Sonnet releases did not have. It also changes some API behavior, so substituting the model identifier requires a migration review.
Operational significance
Measure latency and cost per completed task under the same harness and effort settings. A high-effort Sonnet call may have a different cost advantage from a low-effort call. Check the current pricing and actual token use when making a budget decision.
Use evaluations to decide which routine work Sonnet can handle and when to escalate a case. The application needs acceptance criteria for the result, whatever model produced it.
Distinguish it from nearby terms
- Claude is the overall family; Sonnet is one line within it.
- Opus targets complex work requiring sustained judgment. Haiku targets high-volume work with tighter latency and cost constraints.
- Model routing chooses a model for a request. A model cascade runs another stage when an earlier result fails a defined acceptance rule.
- Latency is measured delay. A line's reputation for speed does not establish the response time of a particular workload or platform.
Check your understanding
Sonnet and Opus receive similar scores on a coding benchmark, but Sonnet needs more retries on your multi-service migrations. Which measurements should determine your routing policy?