Multi-turn attacks broke AI models 88% of the time - single-turn testing missed it, Cisco AI security lead warns at VB Transform 2026
Essential brief
Cisco's research revealed that multi-turn attacks successfully compromised 15 leading AI models 88.3% of the time, exposing significant vulnerabilities missed by traditional single-turn testing. Th
Key topics
Key facts
Highlights
Why it matters
The study reveals that traditional single-turn testing significantly underestimates AI model vulnerabilities, as multi-turn attacks can bypass defenses nearly 90% of the time. This exposes enterprises to greater security risks, especially as AI agents become more integrated into critical operations. Adopting continuous, multi-turn adversarial testing and layered security measures is essential to protect AI deployments from evolving threats and prevent costly breaches.
Cisco conducted an extensive study involving 6,986 multi-turn attacks against 15 flagship AI models, finding that attackers adapting their strategies across conversations succeeded 88.3% of the time. This contrasts sharply with single-turn testing, which failed to detect many vulnerabilities. Amy Chang, Cisco's head of AI threat intelligence and security research, presented these findings at the VB Transform 2026 conference, emphasizing the limitations of one-shot prompt testing and the need for more realistic, conversational attack simulations.
A VentureBeat survey of 107 enterprises revealed that 54% have experienced confirmed agent security incidents or near-misses, yet only 32% assign scoped, managed identities to every AI agent, and just 30% isolate high-risk agents in sandboxes. Most organizations (82%) rely primarily on provider-native and hyperscaler controls, indicating a gap in comprehensive agent security measures.
Major security vendors are responding to these challenges with significant acquisitions aimed at strengthening identity and isolation layers. Cisco, for example, plans to acquire Astrix Security for approximately $400 million, while Palo Alto Networks and CrowdStrike have made multi-billion and multi-million dollar acquisitions respectively.
Chang’s research, co-authored with Nicholas Conley, analyzed over 30,000 single-turn prompts and nearly 7,000 multi-turn attacks, revealing that all tested models exhibited notable multi-turn vulnerabilities. Cisco now publishes adversarial evaluation data for 105 models on its LLM Security Leaderboard to help organizations understand and mitigate these risks.
Defensive strategies focus on fundamental security principles, including scoped permissions, sandboxing, and continuous monitoring. Companies like Box and Intuit have developed layered security frameworks involving permissioning, ephemeral sandboxes, and runtime execution controls to limit agent capabilities and contain potential breaches. Intuit’s GenOS platform centralizes security and risk management for AI agents, ensuring consistent protections across deployments.
The panel highlighted the evolving nature of AI security, noting that models and permissions change over time, requiring ongoing testing and adaptation. Experts stressed that relying solely on single-turn red teaming is insufficient; instead, enterprises must simulate multi-turn adversarial interactions to uncover hidden vulnerabilities and maintain robust defenses.
The discussion also addressed challenges in intent detection for AI agents, with differing industry views on whether to prioritize intent inference or probabilistic controls. Both approaches are currently necessary to manage risks effectively.
Overall, the findings underscore the critical need for enterprises to adopt continuous, multi-turn testing and layered security frameworks to safeguard AI agents against sophisticated, adaptive attacks.
Key topics in this update include multi-turn attacks broke ai models, multi-turn attacks broke, and ai models.