AI SecurityResearchers Want to Read an AI's Mind Before It Does Something Dangerous
A university team is building tools to watch what happens inside an AI model as it thinks, not just what it says. The goal: catch harmful requests that slip past every other filter.