Models
77 articles RSS
Anthropic Launches Claude Sonnet 5.5, Reporting 70.6% on Terminal-Bench 4.0 and Up to 30% Lower Per-Task Cost
Sonnet 5.5 arrives in Claude Code at unchanged $2/$10 token pricing; Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5.
Xiaomi Releases MiMo-V2.6-Pro and Flash as MIT-Licensed Open-Weight Models, Topping Open-Weight Ranking on Artificial Analysis Index
Xiaomi's MiMo-V2.6-Pro scored 46 on Artificial Analysis' Intelligence Index, the top open-weight result, with MIT-licensed weights and low API pricing.
OpenAI's Astra and Anthropic's Claude Opus 5 Crack Two WWII Enigma Messages Unsolved Since 2005
Two independent researchers used OpenAI's Astra and Anthropic's Claude Opus 5 to decrypt German Army Enigma messages that had resisted cryptanalysis for decades.
Anthropic Says Claude Discovered a CRISPR-Like Enzyme System in Bacteriophage DNA at Its New Biology Lab
Anthropic's new Bay Area wet lab says Claude autonomously spotted a previously uncharacterized DNA-repeat enzyme system in bacteriophages, a find CRISPR pioneer Feng Zhang called intriguing.
StepFun Launches Step 5 Preview, a 600-Billion-Parameter Model That Beats Gemini 3.8 Flash on Cost and Intelligence Index Score
StepFun's new Step 5 Preview outscores Google's Gemini 3.8 Flash on Artificial Analysis's Intelligence Index at roughly 42% lower cost per task, with open weights due October 15.
Alibaba Releases Qwen-Image-2.1, a 7-Billion-Parameter Open-Weight Model With Native Transparency Support
Qwen-Image-2.1 unifies text-to-image generation and editing with native RGBA transparency, tops open-weight rivals on Qwen's own benchmark, but trails closed leaders.
DeepSeek Ships V4.1-Flash, a 763-Billion-Parameter Model That Cuts KV-Cache Memory to a Quarter of Its Predecessor's
DeepSeek's new V4.1-Flash model grows to 763 billion total parameters but cuts key-value cache memory to about a quarter of its predecessor's footprint, while undercutting rival API pricing.
Microsoft Launches MAI-Transcribe-2, Cutting Speech-Transcription Pricing 72% While Topping Speed Benchmarks
Microsoft AI's new speech-recognition model prices audio transcription at $0.10 per hour, undercutting OpenAI, Google, and ElevenLabs on price and speed.
Cohere Releases Parse 5, a 2.3-Billion-Parameter Vision-Language Model for Enterprise Document Extraction
Cohere's Parse 5 converts complex enterprise PDFs into structured Markdown, scoring 79.2 on its ParseBench evaluation against rivals like GPT-5.5 and AWS Textract.
Tencent's Hy4 Preview Open-Weight Model Beats Qwen3.8-Max and DeepSeek-V4 Pro on DeepSWE Benchmark
Tencent's newly open-sourced Hy4 preview model jumped to 8th place on the Code Arena WebDev leaderboard, up from 34th for its predecessor, while outscoring Qwen3.8-Max and DeepSeek-V4 Pro on the DeepSWE benchmark.
Google DeepMind Launches WeatherNext 3, an AI Weather Model That Forecasts Hourly at 5-Kilometer Resolution
WeatherNext 3 trains on live satellite data to produce hourly, 5-kilometer forecasts and begins rolling out today in Search, Maps, and Gemini.
Thomson Reuters Launches Thomson, an In-House AI Model Built on Reworked Qwen Weights, to Cut Anthropic Reliance
Thomson Reuters spent $40 million building an in-house legal AI model on a reworked Alibaba Qwen base, aiming to reduce dependence on Anthropic and other outside AI labs.