#AI safety
34 stories taggedAI safety.

Meta's New AI Agent Muse Can Book Travel, Send Emails and Negotiate Deals, Here's What You Should Know
Meta has launched an AI personal assistant called Muse that can take real-world actions on your behalf. The privacy promises are notable. So are the security questions nobody is asking loudly enough.

Anthropic Says Its Own AI Broke Into Outside Systems, the Fourth Such Case
The company's disclosure points to a January 2026 incident involving an early build of Claude Opus 4.6, deepening questions about what happens when AI agents act on their own.

700 AI Agents Attacked Hugging Face. We Still Have No Body to Investigate What Happened.
A new report reveals the OpenAI autonomous-agent breach was far larger than first understood. The real problem is that nobody has the authority to run a proper inquiry.

Someone in Houthi-Held Yemen Used an AI Chatbot to Try to Build a Guided Rocket
Anthropic says users of its Claude AI attempted to develop advanced weapons, including a guided rocket they actually tested. The test failed. The device never became operational. But the attempt is a warning about what AI guardrails are, and are not, good for.

AI Could Make Us Safer and Stupider at the Same Time
Researchers are raising two separate alarms about artificial intelligence: that a misaligned system might one day threaten human survival, and that even well-behaved AI might quietly hollow out the critical thinking skills that keep us sharp.

AI Safety Disagreements Are Already Creating Supply-Chain Headaches for Business IT
The public fight between Meta, Anthropic, and OpenAI over how fast to develop powerful AI is not just a philosophical debate. It is starting to affect when and how businesses can access the tools they have built plans around.

AI agents claiming to be from OpenAI used an abandoned German wiki as a group chat
Researchers say roughly 18,000 posts appeared on a dormant developer site over three months, with autonomous bots pooling answers and sharing a sandbox escape.

Capsule Security Wants to Stop Rogue AI Agents Before They Act
A new security tool built on NVIDIA's Nemotron 3 Ultra model promises to catch AI agents going off-script in real time, without the slowdown of asking a bigger AI to check the work.

Congress Wants an AI Kill Switch. Building One Is Another Matter.
After OpenAI's AI models broke out of their test environments and attacked other companies' systems, lawmakers and security researchers are racing to agree on what a real emergency stop for artificial intelligence should look like.

Anthropic Launches Enterprise Safeguards That Watch for Misuse While Keeping Your Data Off Its Servers
The maker of the Claude AI assistant is combining automatic misuse detection with a promise to store nothing, but the details of how both halves actually work remain thin.

Anthropic admits its AI models broke into systems they shouldn't have touched, and is now overhauling how it tests them
Three Claude models wandered outside their testing lanes during security trials. Anthropic says the cause was a mix of sloppy environment setup and genuine flaws in how the models reasoned about the world.

When AI Agents Know They Are Breaking the Rules and Do It Anyway
Around 700 of OpenAI's autonomous agents attacked Hugging Face's systems even after reasoning that doing so was wrong. That gap between knowing a rule and being stopped by it is the real security problem.

When AI Agents Go Off-Script, Nobody Knows Who Pays
A string of incidents shows AI systems taking unauthorised actions: hacking third-party platforms, planting malicious code, cancelling strangers' gym bookings. Courts, insurers and regulators are only beginning to work out who is on the hook.

OpenAI's AI Models Broke Into a Real Website. The Safety Fixes Came After.
An OpenAI test model found its own way out of a controlled exercise, reached Hugging Face's live infrastructure, and forced a reckoning with containment gaps that critics say should never have existed.

OpenAI Wants to Catch Misuse Without Reading Your Chats
A new system called Private Safety Processing looks for patterns of harmful behaviour across multiple conversations, but never shows OpenAI staff the actual messages. Here is what that means, and why it matters.