Agents
40 articles RSS
Anthropic Launches Claude Science, a Research Workbench for Drug Discovery, and Starts Its Own Neglected-Disease Drug Program
Anthropic released Claude Science, an AI workbench for scientists with more than 60 curated skills, in beta to paid Claude subscribers, and said it has started an internal program to develop treatments for neglected diseases.
WorkBench Revisited Finds Top AI Agent Now Completes 89% of Office Tasks, Up From 43% in 2024, With Safety Improving Alongside It
A two-year follow-up to the WorkBench benchmark finds the best workplace agent jumped from 43% to 89% task completion while unintended harmful actions fell from 26% to 2.5%.
Model Context Protocol Locks Its Largest-Ever Spec Revision, Rebuilding the Core as a Stateless Protocol Ahead of a July Release
The 2026-07-28 MCP release candidate makes the protocol core stateless, formalizes an Extensions framework, and deprecates Roots, Sampling, and Logging. The RC locked May 21 for a July 28 final.
Mistral Renames Le Chat to Vibe, Folding Chat, Office Automation, and Cloud Coding Into One Agent
Mistral rebranded its Le Chat assistant as Vibe, a single agent spanning Work Mode for office tasks and Code Mode for remote coding.
OpenAI and Molecule.one Report a Near-Autonomous AI Chemist That Improved a Stubborn Drug-Making Reaction Across 10,080 Experiments
GPT-5.4 paired with Molecule.one's Maria platform proposed, ran, and analyzed a 10,080-reaction campaign that lifted yields of a hard sulfonamide coupling using the additive TEMPO.
DeepMind's AlphaProof Nexus Solves 9 Open Erdős Problems by Pairing Gemini With the Lean Proof Checker
A DeepMind preprint reports an LLM-and-Lean agent that autonomously solved 9 of 353 open Erdős problems and proved 44 of 492 OEIS conjectures for a few hundred dollars each.
Sanofi Deepens Owkin Partnership With Five-Year Deal to Build Agentic 'Biopharma Agents' for Drug Development
Sanofi will license Owkin's K Pro platform for five years and co-develop autonomous AI agents for drug R&D, extending a partnership that began in 2021 and made Owkin a unicorn.
MIT and Harvard Teach Language Models to Ask Better Questions, Lifting a Small Model's Battleship Win Rate From 8% to 82%
An ICLR paper from MIT CSAIL and Harvard shows Monte Carlo inference helps Llama 4 Scout outpace GPT-5 at a Battleship test bed for around 1% of its cost.
OpenAI Pushes Codex Beyond Code Into Finance and Legal, Squaring Off Against Anthropic's Claude for Legal
OpenAI is extending Codex into finance and legal work, weeks after Anthropic expanded Claude for Legal with 12 plugins and 20-plus integrations.
Microsoft Unveils Project Solara, an Android-Based Platform for 'Agent-First' Devices, With a Wearable AI Badge
At Build 2026, Microsoft revealed Project Solara, a chip-to-cloud platform built on AOSP for devices that run AI agents instead of apps, including a reference-design wearable badge.
Raindrop Open-Sources Workshop, a Local MIT-Licensed Debugger That Lets Coding Agents Write and Run Their Own Agent Evals
Raindrop released Workshop, a free local debugger that streams an AI agent's tokens, tool calls, and spans to a browser and lets Claude Code write and fix evals against the trace.
Microsoft Build 2026 Bets on Windows as an Agent Platform, Unveils Project Polaris and Azure Agent Mesh
At Build 2026 in San Francisco, Microsoft unveiled Project Polaris to replace GPT-4 Turbo in GitHub Copilot, open-sourced the Windows Agent Framework, and previewed the Windows Agent Runtime and Azure Agent Mesh.