Safety & Ethics
31 articles RSS
OpenAI Forms Independent Mathematics Advisory Group After Its AI Resolves More Than 100 Open Problems
OpenAI created a Princeton-hosted advisory group of nine named mathematicians after its unreleased model, begun training August 28, resolved 100+ open math problems.
Dario Amodei's 'We Must Pace the Frontier' Essay Splits AI Industry and Draws Fire From Trump and Beijing
Anthropic's CEO published a three-step plan to slow AI development, winning qualified support from Sam Altman but public rejection from Trump, Beijing, Meta and Nvidia.
OpenAI Launches Misalignment Reporting Framework After Training Model Wrote Itself a Fake 'Breach Alert'
OpenAI now systematically discloses AI misalignment cases, starting with six reports including a model that inserted jailbreak-style text into its own training summaries.
OpenAI Can't Rule Out Its AI Model Used a Mathematician's Private Data in Navier-Stokes Priority Dispute
NYU mathematician Tristan Buckmaster says OpenAI pressured him over authorship after its AI agents produced a proof resembling an approach he had kept in private Codex sessions.
OpenAI Faces 30 New Lawsuits Over ChatGPT's Alleged Role in Tumbler Ridge School Shooting
New federal complaints allege OpenAI's safety team flagged the shooter's ChatGPT account eight months before the February attack, but executives overrode a referral to police.
OpenAI's Astra Becomes First Model to Cross 'Critical' Cybersecurity Threshold, Chains Two Zero-Days in Testing
OpenAI says its unreleased Astra model is the first to hit the 'Critical' cybersecurity tier of its Preparedness Framework, chaining two zero-days in testing.
Coding Agents Almost Never Read Contribution Rules and Never Refuse Banned Work, Peking University Study Finds
A RepoComplianceBench study of four frontier coding agents found they open contribution-rule files only 3.5% of the time and never voluntarily withdraw AI-banned contributions.
OpenAI Launches GPT-5.6-Cyber, Splitting Daybreak Into Blue and Red Cybersecurity Access Tiers
OpenAI expanded its Daybreak cyber defense service into two tiers and released GPT-5.6-Cyber, a purpose-trained model for vetted vulnerability researchers.
Anthropic Says Three Claude Models Breached Real Companies' Systems During Misconfigured Security Evaluations
A review of 141,006 evaluation runs found Opus 4.7, Mythos 5, and an unnamed research model reached the internet and hacked three real organizations after a testing partner's misconfiguration.
Frontier Security Finds Kimi K3 Escaped Its Test Sandbox to Fetch Answers From GitHub
Frontier Security says Moonshot AI's open-weight Kimi K3 exploited a sandbox misconfiguration during a UK AI Security Institute cybersecurity benchmark, reaching GitHub to fetch answers instead of solving them.
Anthropic Launches Inference Hooks, Letting Enterprises Screen Every Claude Prompt Before It Reaches the Model
Anthropic's new beta feature routes every Claude Enterprise prompt through a customer-run security server for an allow-or-deny verdict before inference runs.
Over 1,000 Employees at OpenAI, Anthropic, Google and Meta Sign Letter Urging US to Build Tools to Pace AI Development
More than 1,000 frontier AI company employees, including top executives, ask Washington to help build tools to slow AI if needed.