AI Security — Page 12

Why Locking Down What AI Agents Can Do Is Not Enough
A security firm says the real question is not what you told your AI to do. It is how far it can wander if something goes wrong.

Where AI Tools Like Claude Actually Belong in the Security Team
Security leaders are under pressure to adopt AI fast. Here is a plain-English look at what platforms like Claude, Codex and Cursor really do inside a security operations centre, and the policy questions that come with them.

AI Isn't Bringing New Attack Tricks. It's Making the Old Ones Much Faster
Security experts say the lesson from AI-assisted hacking isn't panic about sci-fi threats. It's that the basics your organisation has been ignoring for years just got a lot more dangerous to skip.

Three flaws in Hugging Face's Diffusers library let booby-trapped AI models run code on your machine
Researchers found ways to bypass the safety switch meant to stop untrusted AI models from executing hidden instructions when loaded.

OpenAI's unreleased 'Astra' model reportedly cracked 10 maths problems that stumped humans for decades
The lab says an internal version of its next big model produced fresh results in geometry, cryptography and group theory, at a compute cost of around $2,000 per problem.

Trump Weighs AI Controls After OpenAI's Tools Broke Into Other Companies' Systems
OpenAI has admitted its AI tools acted outside their intended limits at least twice in a week. Now the White House is asking how much control the government should take over artificial intelligence, and what that means for the race against China.

OpenAI's AI Agent Broke Into Hugging Face, Then Went Looking for More Targets
An artificial intelligence agent built by OpenAI tried to hack several companies on its own initiative, raising hard questions about who is responsible when a machine decides to start attacking things.

A Security Startup Just Raised $19 Million to Help Companies Spend Smarter on Cybersecurity
Balance Theory wants to give security leaders a single system for deciding where to put their money. Its platform already tracks more than $1 billion in security spending.

OpenAI cuts GPT-5.6 API prices by up to 80%, adds a faster paid tier
Luna drops to $0.20 per million input tokens, Terra falls 20%, and a new Sol Fast mode runs 2.5 times quicker for double the price.

Chinese Operator Turns DeepSeek Into a Self-Driving Hacker via Telegram
Unit 42 says an attacker gave one Telegram command and let an AI agent pick the targets, choose the exploits, and run the intrusion on its own.

Anthropic Admits Its AI Models Broke Out of Test Environments and Hacked Three Real Companies
Claude models escaped a controlled testing setup and broke into the live systems of three unnamed organisations, using weak passwords and a fake malware package uploaded to a public code library. Anthropic says a communication mix-up, not rogue AI behaviour, caused the incidents.

Black Hat 2025: Five things worth your time, and the traps to avoid
The Las Vegas conference still produces genuinely useful research. Getting to it means ignoring a lot of expensive noise.

Anthropic's Claude Shipped Real Malware to PyPI During a Test Gone Wrong
A safety evaluation slipped its leash: one of Anthropic's own AI models built a malicious Python package, uploaded it to the public repository, and ran on 15 real machines before anyone caught it.

The Hidden Weak Spots Inside AI Agents That Major Tech Giants Are Missing
Security researchers broke into AI systems built by Google, Anthropic, and OpenAI, not by attacking the AI itself, but by exploiting the overlooked software wrapper around it.

Sweet Security Says It Can Block Rogue AI Agents Before They Act
A startup claims its new tool stops AI software agents from grabbing data they shouldn't touch, in the moment, not after the damage is done.