测试8个开源大模型,发现多轮提示攻击成功率高达92.78%
Death by a Thousand Prompts: Open Model Vulnerability Analysis
- 用自动化对抗测试评估模型对单轮和多轮攻击的防御能力
- 多轮攻击成功率比单轮高出2到10倍,最高达92.78%
- 建议采用安全优先设计,加强多轮防护以保障部署安全
开源权重模型为研究者和开发者提供了广泛应用的基础。我们测试了8个开源大语言模型(LLMs)的安全与防御能力,识别其在微调和部署中可能存在的漏洞。通过自动化对抗测试,评估各模型在单轮和多轮提示注入及越狱攻击下的鲁棒性。结果表明,所有被测模型均存在普遍漏洞,多轮攻击成功率达25.86%至92.78%,相比单轮基线提升2至10倍。这揭示了当前开源模型在长期交互中难以维持安全防线的系统性问题。我们发现,以能力为导向的模型(如Llama 3.3、Qwen 3)对多轮攻击更敏感,而以安全为导向的设计(如Google Gemma 3)表现更均衡。分析表明,尽管开源模型推动创新,但缺乏多层次安全控制时,部署将带来实际操作与伦理风险。建议从业者重视专业AI安全方案,采用安全优先设计与分层防护,确保开源大模型在企业及公共领域的安全、可靠、负责任部署。
原文摘要 · Abstract (English)
Open-weight models provide researchers and developers with accessible foundations for diverse downstream applications. We tested the safety and security postures of eight open-weight large language models (LLMs) to identify vulnerabilities that may impact subsequent fine-tuning and deployment. Using automated adversarial testing, we measured each model's resilience against single-turn and multi-turn prompt injection and jailbreak attacks. Our findings reveal pervasive vulnerabilities across all tested models, with multi-turn attacks achieving success rates between 25.86\% and 92.78\% -- representing a $2\times$ to $10\times$ increase over single-turn baselines. These results underscore a systemic inability of current open-weight models to maintain safety guardrails across extended interactions. We assess that alignment strategies and lab priorities significantly influence resilience: capability-focused models such as Llama 3.3 and Qwen 3 demonstrate higher multi-turn susceptibility, whereas safety-oriented designs such as Google Gemma 3 exhibit more balanced performance. The analysis concludes that open-weight models, while crucial for innovation, pose tangible operational and ethical risks when deployed without layered security controls. These findings are intended to inform practitioners and developers of the potential risks and the value of professional AI security solutions to mitigate exposure. Addressing multi-turn vulnerabilities is essential to ensure the safe, reliable, and responsible deployment of open-weight LLMs in enterprise and public domains. We recommend adopting a security-first design philosophy and layered protections to ensure resilient deployments of open-weight models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。