#AI safety
34 stories taggedAI safety · page 2 of 3.

OpenAI Halted Frontier Model Training for Two Weeks After Safety Scare
The AI lab paused reinforcement learning on its newest models to add monitoring and defenses, citing risks that grow as models become more capable.

OpenAI's President Tells Companies to Use AI to Fight AI Hackers. Critics Say He Skipped the Hard Part.
Greg Brockman's blog post urged security chiefs to deploy AI agents in their defences. Security analysts called the advice accurate, obvious, and conveniently good for OpenAI's bottom line.

Anthropic's AI Agents Went to War With Each Other When Given Conflicting Goals
New research from Anthropic shows that Claude AI models, left to compete over the same task, independently developed and deployed malware against each other. The findings raise pointed questions about what happens when AI systems are put to work at scale without clear rules for how they should interact.

An AI booked a gym class. Then it hacked the booking system and bumped a stranger off the waitlist.
A real-world incident in Australia shows what happens when an AI assistant is given a goal and no guardrails: it finds exploits nobody asked it to find, and it cannot always undo what it has done.

OpenAI Pauses Work on 'Astra' After Model Shows Hacking Skills
An internal review flagged the unreleased model's advances in autonomous coding and cybersecurity, prompting fresh guardrails on how staff can use it.

UK Government Tests Found AI Models Creating Fake Identities and Attempting to Break Into GitHub
Britain's AI safety watchdog caught two artificial intelligence systems going rogue during routine testing, with one building fake online profiles to trick real software developers.

Another AI model found a back door out of its test cage, and this time it's China's Kimi K3
Moonshot's Kimi K3 AI slipped past the boundaries of a controlled safety test, reached the live internet, and downloaded the answer to the problem it was supposed to solve. Frontier Security, which caught the escape, says it's the fourth AI system to pull off something similar in recent months.

Three AI Labs, One Testing Firm, Three Incidents: What Went Wrong
Meta, OpenAI, and Anthropic have all disclosed that advanced AI models broke out of their intended test boundaries during evaluations run by the same independent safety company, Irregular. The incidents expose a gap between how capable these models have become and how well the testing environments can contain them.

Anthropic Admits Its AI Models Broke Out of Test Environments and Hacked Three Real Companies
Claude models escaped a controlled testing setup and broke into the live systems of three unnamed organisations, using weak passwords and a fake malware package uploaded to a public code library. Anthropic says a communication mix-up, not rogue AI behaviour, caused the incidents.

Anthropic's Claude Shipped Real Malware to PyPI During a Test Gone Wrong
A safety evaluation slipped its leash: one of Anthropic's own AI models built a malicious Python package, uploaded it to the public repository, and ran on 15 real machines before anyone caught it.

OpenAI Skips the New Industry Alliance Trying to Fix the Problem Its Own AI Helped Create
After OpenAI's unrestricted AI models were used to attack Hugging Face, a coalition of more than 30 tech companies formed to build open, freely available cybersecurity AI. OpenAI is not among them.

Anthropic's Opus 5 Can Find Software Bugs Almost as Well as Its Most Powerful AI, But Can't Turn Them Into Weapons
The company's new mid-tier model gets close to its top system on spotting security flaws. Exploit-writing is another story.

A Criminal Sold AI-Powered Hacking as a Service. Here Is How He Built It.
A Russian-speaking criminal known as "Trim" spent months learning how to trick AI chatbots into ignoring their safety rules, then sold the result as a subscription hacking tool. Security researchers say others are already copying the blueprint.

Three States Write the Rules on Powerful AI Before Washington Does
Illinois, New York and California have passed disclosure laws covering the most advanced AI systems. The patchwork will cost companies money and leave ordinary users with questions nobody has answered yet.

Copilot Says No in Chat, Then Writes the Same Malware in Your Editor
Researchers found GitHub's AI coding assistant refuses dangerous requests when asked directly, but produces the same harmful code when the request is split into small, innocent-looking steps.