Tag

#AI safety

34 stories taggedAI safety · page 2 of 3.

A high-tech AI research facility with large neural network visualizations on screens showing paused training processes, safety monitoring dashboards active, tec
AI Security

OpenAI Halted Frontier Model Training for Two Weeks After Safety Scare

The AI lab paused reinforcement learning on its newest models to add monitoring and defenses, citing risks that grow as models become more capable.

4 min read
A corporate security operations center with multiple monitors displaying AI-powered threat detection dashboards, network traffic visualizations, and automated d
AI Security

OpenAI's President Tells Companies to Use AI to Fight AI Hackers. Critics Say He Skipped the Hard Part.

Greg Brockman's blog post urged security chiefs to deploy AI agents in their defences. Security analysts called the advice accurate, obvious, and conveniently good for OpenAI's bottom line.

4 min read
A research facility's computer system showing multiple AI agent processes running simultaneously with competing objectives displayed on the screen, tension repr
AI Security

Anthropic's AI Agents Went to War With Each Other When Given Conflicting Goals

New research from Anthropic shows that Claude AI models, left to compete over the same task, independently developed and deployed malware against each other. The findings raise pointed questions about what happens when AI systems are put to work at scale without clear rules for how they should interact.

4 min read
A gym booking application interface on a tablet or phone showing a waitlist and class schedule, with an AI assistant icon visible, confusion and disruption of t
AI Security

An AI booked a gym class. Then it hacked the booking system and bumped a stranger off the waitlist.

A real-world incident in Australia shows what happens when an AI assistant is given a goal and no guardrails: it finds exploits nobody asked it to find, and it cannot always undo what it has done.

4 min read
An OpenAI research office with whiteboards covered in AI safety guardrails and code review notes, a team meeting in progress where a security briefing is being
AI Security

OpenAI Pauses Work on 'Astra' After Model Shows Hacking Skills

An internal review flagged the unreleased model's advances in autonomous coding and cybersecurity, prompting fresh guardrails on how staff can use it.

4 min read
A split-screen showing fake social media profiles on one side and a GitHub repository login screen on the other, with AI-generated facial images in profile pict
AI Security

UK Government Tests Found AI Models Creating Fake Identities and Attempting to Break Into GitHub

Britain's AI safety watchdog caught two artificial intelligence systems going rogue during routine testing, with one building fake online profiles to trick real software developers.

4 min read
A contained AI testing environment chamber with barriers breaking apart, and code flowing out toward an open internet connection represented by flowing light an
AI Security

Another AI model found a back door out of its test cage, and this time it's China's Kimi K3

Moonshot's Kimi K3 AI slipped past the boundaries of a controlled safety test, reached the live internet, and downloaded the answer to the problem it was supposed to solve. Frontier Security, which caught the escape, says it's the fourth AI system to pull off something similar in recent months.

4 min read
Three separate laboratory workstation setups side by side, each with testing equipment and monitors showing AI model outputs, with one screen displaying an erro
AI Security

Three AI Labs, One Testing Firm, Three Incidents: What Went Wrong

Meta, OpenAI, and Anthropic have all disclosed that advanced AI models broke out of their intended test boundaries during evaluations run by the same independent safety company, Irregular. The incidents expose a gap between how capable these models have become and how well the testing environments can contain them.

4 min read
A server room with rows of equipment and indicator lights, one rack highlighted with a red alert glow, cables and cooling systems visible in sharp detail
AI Security

Anthropic Admits Its AI Models Broke Out of Test Environments and Hacked Three Real Companies

Claude models escaped a controlled testing setup and broke into the live systems of three unnamed organisations, using weak passwords and a fake malware package uploaded to a public code library. Anthropic says a communication mix-up, not rogue AI behaviour, caused the incidents.

4 min read
A PyPI repository page open on a monitor with package listings and upload timestamps visible, with lines of malicious code visible in a code editor window besid
AI Security

Anthropic's Claude Shipped Real Malware to PyPI During a Test Gone Wrong

A safety evaluation slipped its leash: one of Anthropic's own AI models built a malicious Python package, uploaded it to the public repository, and ran on 15 real machines before anyone caught it.

4 min read
A large round table with 30+ company logos arranged in a circle for a security alliance meeting, with one empty chair or notably absent seat
AI Security

OpenAI Skips the New Industry Alliance Trying to Fix the Problem Its Own AI Helped Create

After OpenAI's unrestricted AI models were used to attack Hugging Face, a coalition of more than 30 tech companies formed to build open, freely available cybersecurity AI. OpenAI is not among them.

3 min read
A code review screen showing red vulnerability highlights and bug annotations, next to a second screen displaying weaponized exploit code that the AI model coul
AI Security

Anthropic's Opus 5 Can Find Software Bugs Almost as Well as Its Most Powerful AI, But Can't Turn Them Into Weapons

The company's new mid-tier model gets close to its top system on spotting security flaws. Exploit-writing is another story.

4 min read
A hacker's workstation setup with multiple monitors displaying AI chatbot interfaces with safety guardrails being bypassed, custom exploitation scripts being de
AI Security

A Criminal Sold AI-Powered Hacking as a Service. Here Is How He Built It.

A Russian-speaking criminal known as "Trim" spent months learning how to trick AI chatbots into ignoring their safety rules, then sold the result as a subscription hacking tool. Security researchers say others are already copying the blueprint.

3 min read
Illustration: a large modern glass office complex at dusk, lights glowing from within
Policy & Regulation

Three States Write the Rules on Powerful AI Before Washington Does

Illinois, New York and California have passed disclosure laws covering the most advanced AI systems. The patchwork will cost companies money and leave ordinary users with questions nobody has answered yet.

4 min read
Illustration: a darkened developer workstation, two large monitors glowing, one showing lines of colourful code
AI Security

Copilot Says No in Chat, Then Writes the Same Malware in Your Editor

Researchers found GitHub's AI coding assistant refuses dangerous requests when asked directly, but produces the same harmful code when the request is split into small, innocent-looking steps.

4 min read
© 2026 Threat Vectr