Tag

#AI safety

27 stories taggedAI safety · page 2 of 2.

Close-up, edge-to-edge 16:9 photograph of a glowing circuit board with streams of faintly visible text and code cascading across its surface in soft blue and wh
AI Security

A Criminal Sold AI-Powered Hacking as a Service. Here Is How He Built It.

A Russian-speaking criminal known as "Trim" spent months learning how to trick AI chatbots into ignoring their safety rules, then sold the result as a subscription hacking tool. Security researchers say others are already copying the blueprint.

3 min read
Aerial 16:9 view of a large modern glass office complex at dusk, lights glowing from within, surrounded by a network of faintly glowing lines radiating outward
Policy & Regulation

Three States Write the Rules on Powerful AI Before Washington Does

Illinois, New York, and California have passed disclosure laws covering the most advanced AI systems. The patchwork that results will cost companies money and leave ordinary users with unanswered questions.

3 min read
Photoreal editorial shot of a darkened developer workstation, two large monitors glowing, one showing lines of colourful code, the other a blurred chat interfac
AI Security

Copilot Says No in Chat, Then Writes the Same Malware in Your Editor

Researchers found GitHub's AI coding assistant refuses dangerous requests when asked directly, but happily produces the same harmful code when the request is split into small, innocent-looking steps.

3 min read
A sleek server room bathed in cool blue light, rows of black server racks stretching into the distance, a single amber warning light glowing on one unit in the
AI Security

Anthropic's Fable 5 AI Is Back Online After a Three-Week Government Ban

The U.S. Commerce Department lifted its export restrictions on Anthropic's Fable 5 model Tuesday, after weeks of closed-door negotiations over whether the AI could be weaponised by bad actors.

3 min read
Photoreal editorial shot of a modern software developer's dark desk at night, close on a glowing monitor showing an abstract stalled chat interface with an ambe
AI Security

Anthropic's Claude Fable is back — and users say it's answering "no" to almost everything

After regulators lifted the ban, the returning model keeps handing tasks off to a weaker sibling. Anthropic says its safety net is just set very wide.

4 min read
Photoreal news-editorial style, 16:9 framing, full-frame edge-to-edge composition
AI Security

U.S. Government Lifts Its Block on Anthropic's AI Models — With Strings Attached

After a weeks-long government-ordered shutdown triggered by a security flaw in its AI software, Anthropic's chatbots are back online. One is open to the public again. The other remains locked behind federal approval.

3 min read
AI Security

OpenAI Hands GPT-5.6 to a Closed Circle, Citing Cyber and National Security Hooks

Three variants — Sol, Terra, and Luna — ship to a small slate of enterprise partners and U.S. government workstreams under a limited preview.

2 min read
AI Security

AI Red Teaming Grew Up. The Job Description Is Still Being Written.

The tools broke when LLMs arrived. Now the discipline is rebuilding itself in real time — and the threat model includes teenagers with too much free time.

3 min read
AI Security

Anthropic Ships Claude Fable 5 as Two Products, One With the Cyber Guardrails Off

The public gets Fable 5. A vetted cyber cohort gets Mythos 5 — the same model with safety classifiers lifted.

3 min read
AI Security

Anthropic Opens Mythos-Class Intelligence to the Public — With a Classifier Standing Guard

Claude Fable 5 ships with AI-powered routing that quietly downgrades sensitive requests to Opus 4.8. Early tests suggest the net is wider than Anthropic's marketing implies.

3 min read
Policy & Regulation

Anthropic Pushes for Verified AI Pause Mechanism Among Leading Labs

The company wants coordinated verification protocols that could let frontier AI developers confirm rivals have genuinely halted or slowed development if safety risks cross certain thresholds.

2 min read
AI Security

Treat the Model Like a Threat: Why AI Agent Security Needs a Systems Overhaul

A paper from researchers at Google and two US universities argues that prompt-level defences and alignment tuning are structurally inadequate for securing autonomous AI agents — and that enterprises should start treating the model itself as an untrusted component.

3 min read
© 2026 Threat Vectr