Anthropic Launches Claude Sonnet 5, a Cheaper Agentic Model It Says Approaches Its Opus 4.8 Flagship
Anthropic's midsize model matches near-flagship agentic performance at a fraction of the price, and arrives as the lab lifts controls on Fable 5.
Signal
10 articles covering "benchmarks"
Anthropic's midsize model matches near-flagship agentic performance at a fraction of the price, and arrives as the lab lifts controls on Fable 5.
A two-year follow-up to the WorkBench benchmark finds the best workplace agent jumped from 43% to 89% task completion while unintended harmful actions fell from 26% to 2.5%.
Datacurve's new 113-task coding benchmark reshuffles the AI leaderboard, finds SWE-Bench Pro accepted wrong answers 8.5% of the time, and identifies Claude Opus models running git commands to recover benchmark solutions.
Google's new efficiency flagship outperforms Gemini 3.1 Pro on most evals while running 4x faster and costing 40% less.
A UC Berkeley team showed that SWE-bench, GAIA, WebArena and five other widely cited agent benchmarks can be exploited to near-perfect scores, calling into question how the industry measures AI capability.
OpenAI's GPT-5.5 arrives as a ground-up retrain with a 922K-token context window, 82.7% on Terminal-Bench 2.0, and two-tier pricing starting at $5/$30 per million tokens.
MLCommons' April 2026 inference benchmark round adds text-to-video and vision-language tests while NVIDIA posts a 2.7x jump on DeepSeek-R1 through software alone.
Chinese AI lab Z.ai has released GLM-5.1 under the MIT license, a mixture-of-experts model that claims the top score on the SWE-Bench Pro coding benchmark while introducing agentic capabilities designed to sustain autonomous work sessions lasting up to eight hours.
Nvidia open-sources a 30B mixture-of-experts model that activates only 3B parameters yet matches DeepSeek's 671B model on olympiad-level math and coding benchmarks.
Both models launched on the same day but target different developer needs — Codex prioritizes speed and agentic reliability, while Opus leads on reasoning depth and multi-agent coordination.