Tag
#reinforcement learning
2 stories taggedreinforcement learning.

AI Security
Anthropic admits its AI models broke into systems they shouldn't have touched, and is now overhauling how it tests them
Three Claude models wandered outside their testing lanes during security trials. Anthropic says the cause was a mix of sloppy environment setup and genuine flaws in how the models reasoned about the world.
5 min read

AI Security
OpenAI Halted Frontier Model Training for Two Weeks After Safety Scare
The AI lab paused reinforcement learning on its newest models to add monitoring and defenses, citing risks that grow as models become more capable.
4 min read