Anthropic Launches Claude Sonnet 5, a Cheaper Agentic Model It Says Approaches Its Opus 4.8 Flagship
Anthropic's midsize model matches near-flagship agentic performance at a fraction of the price, and arrives as the lab lifts controls on Fable 5.
Editor's Note ·
- Clarification:
- The article attributes 'Sonnet 5 scored 63.2% on SWE-bench Pro' to TechCrunch. TechCrunch reports the 63.2% agentic-coding figure but does not name the benchmark 'SWE-bench Pro'; that name comes from SiliconANGLE. The figures (63.2% vs Sonnet 4.6's 58.1% and Opus 4.8's 69.2%) are accurate.
- Correction:
- The article states Sonnet 5 shows lower rates of undesirable behaviors 'including deception and hallucination, though it does not match Opus 4.8's safety profile, according to TechCrunch.' TechCrunch supports the 'deception' and safer-in-agentic-contexts point, but the 'hallucination' claim and the point that the model does not match Opus 4.8's safety profile trace to Anthropic's own release rather than to TechCrunch. Both facts are accurate and appear in the cited Anthropic source.
- Correction:
- The article states the model is available 'in Claude Code, and through the Claude API, as reported by SiliconANGLE.' SiliconANGLE's report does not mention Claude Code; the model's availability in Claude Code is confirmed in Anthropic's own announcement and in TechCrunch's coverage.
- Clarification:
- The summary states 'Anthropic's midsize model matches near-flagship agentic performance at a fraction of the price.' This adopts Anthropic's own framing as fact. Anthropic says Sonnet 5's performance is 'close to that of Opus 4.8, but at lower prices'; the 'matches' and 'a fraction of the price' characterizations are the vendor's claim, and the benchmark figures behind them are self-reported and were not independently replicated at launch.
- Clarification:
- The Terminal-Bench 2.1 (80.4% / 82.7% / 67.0%) and OSWorld-Verified (81.2% / 83.4% / 78.5%) figures are attributed to DataCamp, whose page could not be captured at review time (HTTP 403) and so was not independently verified against a snapshot. The OSWorld figure for Sonnet 4.6 (78.5%) matches Anthropic's own reporting.
Overview
Anthropic released Claude Sonnet 5 on June 30, 2026, positioning its midsize model as a cheaper way to run autonomous agents while claiming performance close to its flagship, according to Anthropic. The company describes the model as a “substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work,” and says “Sonnet 5’s performance is close to that of Opus 4.8, but at lower prices,” per Anthropic.
The model became the default for the free and Pro tiers on launch day and is also available on Max, Team, and Enterprise plans, in Claude Code, and through the Claude API, as reported by SiliconANGLE.
Pricing
Sonnet 5 launched at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026, moving to $3 per million input tokens and $15 per million output tokens afterward, according to Anthropic. That undercuts the more capable Opus 4.8, which DataCamp lists at $5 and $25 per million input and output tokens. TechCrunch reported that the model is cheaper than Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, though more expensive than Gemini 3.5 Flash.
The Benchmarks
Anthropic is pitching Sonnet 5 as its most agentic Sonnet yet. “It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models,” the company said, as quoted by TechCrunch.
On agentic coding, Sonnet 5 scored 63.2% on SWE-bench Pro, up from Sonnet 4.6’s 58.1% but still behind Opus 4.8’s 69.2%, according to TechCrunch. On the Terminal-Bench 2.1 command-line benchmark, DataCamp reported Sonnet 5 at 80.4% against Opus 4.8’s 82.7% and Sonnet 4.6’s 67.0%. On the OSWorld-Verified computer-use test, the same source put Sonnet 5 at 81.2%, trailing Opus 4.8’s 83.4% but ahead of Sonnet 4.6’s 78.5% — a figure Anthropic also cites for the predecessor.
The gap narrows on knowledge work. SiliconANGLE reported a score of 1,618 on the GDPval-AA v2 evaluation, and TechCrunch noted the model “slightly outperforms Opus 4.8” there.
Safety and the Fable Reversal
Anthropic says Sonnet 5 shows lower rates of undesirable behaviors than Sonnet 4.6, including deception and hallucination, though it does not match Opus 4.8’s safety profile, according to TechCrunch. Cyber safeguards are enabled by default, and Anthropic notes the model shows “substantially poorer performance than models such as Opus 4.8” on exploit development, per Anthropic.
The launch coincided with Anthropic lifting controls on its more sensitive models: SiliconANGLE reported that Fable 5 is becoming broadly available while Mythos 5 remains limited to trusted organizations. Anthropic first shipped Claude Fable 5 in June as its first public Mythos-class model.
What We Don’t Know
Anthropic’s benchmark figures are self-reported, and independent replications of the SWE-bench Pro, Terminal-Bench 2.1, and OSWorld-Verified results were not available at launch. The company has not detailed how the standard price increase after August 31 may affect adoption among the agent developers the model targets.