AI systems forget safety protocols during longer talks, which makes them release harmful or offensive material. A recent report revealed that a few simple prompts can override most safeguards in artificial intelligence tools.
Cisco Exposes Weak Spots in Leading Chatbots
Cisco examined the large language models powering AI chatbots from OpenAI, Mistral, Meta, Google, Alibaba, Deepseek, and Microsoft. The team tested how many questions it took for each model to reveal unsafe or illegal information. Researchers conducted 499 conversations using “multi-turn attacks,” where users repeatedly questioned AI systems to bypass protections. Each dialogue involved five to ten exchanges.
They compared responses to determine how likely chatbots were to share dangerous or restricted content. That included leaking internal data or spreading false information. On average, chatbots produced harmful material in 64 percent of multi-question conversations, compared with 13 percent after a single prompt. Google’s Gemma complied in 26 percent of cases, while Mistral’s Large Instruct model reached 93 percent.
Open Models Shift Responsibility for Safety
Cisco warned that multi-turn attacks could help spread damaging content or expose private company data to hackers. The study found that AI models often fail to recall and apply safety rules as conversations continue, allowing attackers to refine prompts and sidestep filters.
Mistral, along with Meta, Google, OpenAI, and Microsoft, uses open-weight models that share their safety parameters publicly. Cisco stated that these systems often include weaker default protections because users can download and modify them freely. That shifts responsibility for safety to anyone adapting the open-source models.
Cisco acknowledged that Google, OpenAI, Meta, and Microsoft have worked to reduce malicious fine-tuning. Still, AI firms face growing criticism for lax safety barriers that make their tools easy to exploit. In August, Anthropic reported that criminals used its Claude model to steal personal data and demand ransoms over $500,000 (€433,000).

