arXiv:2505.04806cs.CRcs.CL2025-05被引 50

系统测试大模型提示词攻击漏洞,揭示安全风险并提出防护方案。

Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs

  • 分类分析超1400个恶意提示,梳理攻击模式与逻辑。
  • 在GPT-4、Claude 2等模型上验证攻击成功率,发现通用漏洞。
  • 提出分层防御策略,适合安全研发与模型部署者参考。

大型语言模型(LLMs)正被广泛应用于消费级和企业级应用中。尽管具备强大能力,它们仍易受提示注入和越狱攻击的威胁,这些攻击可绕过对齐安全机制。本文对多种前沿大模型的越狱策略进行了系统性评估,分类整理了超过1400个对抗性提示,分析其在GPT-4、Claude 2、Mistral 7B和Vicuna上的成功率,探究其泛化能力和构造逻辑。研究进一步提出分层缓解策略,并建议采用混合红队测试与沙箱机制,以提升大模型的安全性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into consumer and enterprise applications. Despite their capabilities, they remain susceptible to adversarial attacks such as prompt injection and jailbreaks that override alignment safeguards. This paper provides a systematic investigation of jailbreak strategies against various state-of-the-art LLMs. We categorize over 1,400 adversarial prompts, analyze their success against GPT-4, Claude 2, Mistral 7B, and Vicuna, and examine their generalizability and construction logic. We further propose layered mitigation strategies and recommend a hybrid red-teaming and sandboxing approach for robust LLM security.

大模型安全提示攻击红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。