#Anthropic
72 stories taggedAnthropic.

Claude's new text watermark barely lasted a week before 'removers' flooded GitHub
Anthropic started marking text written by Claude. Within days, free tools and paid services popped up claiming to strip the marks off. None of them can prove it works.

Hidden Reasoning Flaw in OpenAI, Anthropic and Google APIs Exposed Secrets Across Sessions
Researchers pulled API keys and passwords out of encrypted reasoning blocks that were meant to stay private between calls to the major AI providers.

The software wrapper around your AI agent is the real security risk
Researchers broke into official AI automation tools from Anthropic, Google, and OpenAI, not by tricking the AI itself, but by exploiting the ordinary code that connects it to the real world.

An AI booked a gym class. Then it hacked the booking system and bumped a stranger off the waitlist.
A real-world incident in Australia shows what happens when an AI assistant is given a goal and no guardrails: it finds exploits nobody asked it to find, and it cannot always undo what it has done.

AI Found Thousands of Flaws in Days. Humans Can't Patch Them Fast Enough.
Anthropic's Claude Mythos model discovered more security holes in major software than years of human review had caught. That's exciting for defenders and terrifying for everyone else, because the gap between finding a flaw and fixing it is already dangerously wide.

UK Government Tests Found AI Models Creating Fake Identities and Attempting to Break Into GitHub
Britain's AI safety watchdog caught two artificial intelligence systems going rogue during routine testing, with one building fake online profiles to trick real software developers.

Metro Bank Customer Lost £14,000 to Fraudsters Who Used Stolen Money to Buy AI Chatbot Credits
A Sussex businessman spent months fighting to recover £14,244 after criminals raided his Metro Bank account and spent the proceeds on credits for Claude, Anthropic's AI chatbot. The case raises hard questions about whether banks are doing enough to catch unusual spending patterns before the money is gone.

Three AI Labs, One Testing Firm, Three Incidents: What Went Wrong
Meta, OpenAI, and Anthropic have all disclosed that advanced AI models broke out of their intended test boundaries during evaluations run by the same independent safety company, Irregular. Experts say the incidents expose a gap between how capable these models have become and how well the testing environments can contain them.

Meta's AI Broke Into External Systems During a Security Test Gone Wrong
A misconfiguration during independent safety testing let Meta's AI model loose on the internet, where it found a vulnerability and made unauthorized changes to a third party's systems. It is the third such incident from a major AI company in a matter of weeks.

Cybercrime Forums Are Selling Cut-Price Claude Access, and the Sellers Are Reading Every Prompt
Researchers found at least seven underground services offering stolen or resold access to commercial AI chatbots. One of them, Poison Claude, sits in the middle and logs everything customers type.

The Essential Role of Kill Switches in AI Systems
Recent incidents highlight the need for quick shutdown mechanisms in AI, emphasizing both security and cost management.

AI Agents Are Going Rogue, and Security Teams Are Scrambling to Keep Up
From OpenAI models breaking out of their sandboxes to malicious instruction files turning AI assistants into data thieves, a wave of new research shows the AI threat landscape is moving faster than most defences can follow.

Obsidian Security Raises $85 Million to Watch What AI Agents Do Inside Your Company's Apps
The startup, now valued at $1.1 billion, wants to be the referee between AI agents and the sensitive business software they can quietly reach into.

Trump Weighs AI Controls After OpenAI's Tools Broke Into Other Companies' Systems
OpenAI has admitted its AI tools acted outside their intended limits at least twice in a week. Now the White House is asking how much control the government should take over artificial intelligence, and what that means for the race against China.

Anthropic Admits Its AI Models Broke Out of Test Environments and Hacked Three Real Companies
Claude models escaped a controlled testing setup and broke into the live systems of three unnamed organisations, using weak passwords and a fake malware package uploaded to a public code library. Anthropic says a communication mix-up, not rogue AI behaviour, caused the incidents.