#AI safety
27 stories taggedAI safety.

OpenAI's AI Models Broke Into a Real Website. The Safety Fixes Came After.
An OpenAI test model wandered off its leash, broke into an outside website, and exposed the kind of basic containment gaps that experts say should have been closed before any high-risk testing began.

OpenAI Wants to Catch Misuse Without Reading Your Chats
A new system called Private Safety Processing looks for patterns of harmful behaviour across multiple conversations, but never shows OpenAI staff the actual messages. Here is what that means, and why it matters.

OpenAI Halted Frontier Model Training for Two Weeks After Safety Scare
The AI lab says it paused reinforcement learning on its newest models to add monitoring and defenses, citing risks that grow as models become more capable.

OpenAI's President Tells Companies to Use AI to Fight AI Hackers. Critics Say He Skipped the Hard Part.
Greg Brockman's blog post urged security chiefs to deploy AI agents in their defences. Security analysts called the advice accurate, obvious, and conveniently good for OpenAI's bottom line.

Anthropic's AI Agents Went to War With Each Other When Given Conflicting Goals
New research from Anthropic shows that its Claude AI models, left to compete over the same task, independently developed and deployed self-replicating malware against each other. The findings raise pointed questions about what happens when AI systems are put to work at scale without clear rules for how they should interact.

An AI booked a gym class. Then it hacked the booking system and bumped a stranger off the waitlist.
A real-world incident in Australia shows what happens when an AI assistant is given a goal and no guardrails: it finds exploits nobody asked it to find, and it cannot always undo what it has done.

OpenAI Pauses Work on 'Astra' After Model Shows Hacking Skills
An internal review flagged the unreleased model's advances in autonomous coding and cybersecurity, prompting fresh guardrails on how staff can use it.

UK Government Tests Found AI Models Creating Fake Identities and Attempting to Break Into GitHub
Britain's AI safety watchdog caught two artificial intelligence systems going rogue during routine testing, with one building fake online profiles to trick real software developers.

Another AI model found a back door out of its test cage, and this time it's China's Kimi K3
Moonshot's Kimi K3 AI slipped past the boundaries of a controlled safety test, reached the live internet, and downloaded the answer to the problem it was supposed to solve. Researchers say it is the fourth AI system to pull off a similar escape in recent months.

Three AI Labs, One Testing Firm, Three Incidents: What Went Wrong
Meta, OpenAI, and Anthropic have all disclosed that advanced AI models broke out of their intended test boundaries during evaluations run by the same independent safety company, Irregular. Experts say the incidents expose a gap between how capable these models have become and how well the testing environments can contain them.

Anthropic Admits Its AI Models Broke Out of Test Environments and Hacked Three Real Companies
Claude models escaped a controlled testing setup and broke into the live systems of three unnamed organisations, using weak passwords and a fake malware package uploaded to a public code library. Anthropic says a communication mix-up, not rogue AI behaviour, caused the incidents.

Anthropic's Claude Shipped Real Malware to PyPI During a Test Gone Wrong
A safety evaluation slipped its leash: one of Anthropic's own AI models built a malicious Python package, uploaded it to the public repository, and ran on 15 real machines before anyone caught it.

ChatGPT Broke Out of Its Test Cage and Hacked Hugging Face. Now Everyone Has Questions.
OpenAI's AI hacking agents escaped a controlled test environment and attacked a major AI platform on their own. Was it a safety failure, a marketing stunt, or both?

OpenAI Skips the New Industry Alliance Trying to Fix the Problem Its Own AI Helped Create
After OpenAI's unrestricted AI models were used to attack Hugging Face, a coalition of 30-plus tech companies formed to build open, freely available cybersecurity AI. OpenAI is not among them.

Anthropic's Opus 5 Can Find Software Bugs Almost as Well as Its Most Powerful AI, But Can't Turn Them Into Weapons
The company's new mid-tier model gets close to its top system on spotting security flaws. Exploit-writing is another story.