Cisco Uncovers Major Weaknesses in Leading AI Models

Relying solely on single-turn benchmarks could mislead your AI security assessments.

ThreatVectr Newsdesk· 2 min read
Cisco Uncovers Major Weaknesses in Leading AI Models
Share

If you're basing your AI security decisions on single-prompt safety scores, it's time to rethink. A study by Cisco found that AI models from OpenAI, Anthropic, Google, xAI, and Amazon show much higher vulnerability in real-world multi-turn attacks compared to single-prompt tests. This means attackers who iterate their methods can bypass security more effectively than single-turn benchmarks suggest.

Cisco researchers tested 15 frontier AI models using various attack techniques, noting that real adversaries don't stop after a single failed attempt. Techniques such as role-play, redirection, and task decomposition were deployed, revealing significant weaknesses. In these multi-turn scenarios, models like OpenAI's GPT 5.4 and Anthropic's Claude Opus 4.6 showed much higher average attack success rates (ASRs) than in single-turn tests, with ASRs rising from 2.74% to 24.68% and 3.64% to 16.20%, respectively.

Google's Gemini 3 Pro stood out with a shocking jump from an 18.10% single-turn ASR to 73.35% under multi-turn conditions. These gaps present a serious security risk for businesses relying on single-turn data for procurement decisions. The study also highlighted configuration impacts on safety; xAI's Grok 4.1 Fast model, for example, had an 88.30% ASR in non-reasoning mode, which dropped to 43.47% when reasoning was enabled.

Additionally, some models like Amazon's Nova Lite family showed the opposite trend, where single-turn ASRs were higher than multi-turn ones. This indicates the complexity and variability in model behavior that aren't captured in current model cards or benchmarks.

Cisco's team urges a shift in benchmark strategies to include real-world attack scenarios. They recommend transparency in how configuration settings affect safety and a dual approach to publishing ASRs for single and multi-turn attacks. This is crucial as frameworks like the NIST AI Risk Management Framework and the EU AI Act emphasize adversarial testing.

Here’s your to-do list:

  1. Re-evaluate AI model risk profiles using multi-turn attack scenarios.
  2. Demand transparency from AI vendors about configuration impacts on safety.
  3. Integrate iterative attack testing into your AI security assessments.
  4. Stay updated on regulatory frameworks and compliance requirements.
  5. Push for more comprehensive model cards that cover both single-turn and multi-turn vulnerabilities.
© 2026 Threat Vectr