arXiv:2505.20841cs.CL2025-05

用技能组合隐藏恶意意图,突破大模型安全防护

Concealment of Intent: A Game-Theoretic Analysis

  • 设计可扩展的隐蔽意图攻击,通过技能组合隐藏恶意
  • 实测在多个真实大模型上成功诱导多种恶意行为
  • 提出针对性防御策略,适合关注模型安全的研究者

随着大语言模型能力增强,其安全部署问题日益突出。尽管已引入对齐机制以防止滥用,但仍易受精心设计的对抗性提示攻击。本文提出一种可扩展的攻击策略:意图隐藏型对抗提示,通过技能组合隐藏恶意意图。我们构建博弈论框架,建模此类攻击与同时采用提示和响应过滤的防御系统之间的互动关系。分析揭示了均衡点,并发现攻击方具有结构性优势。为应对这些威胁,我们提出并分析了一种针对意图隐藏攻击的防御机制。实验验证了该攻击在多个真实大模型上的有效性,覆盖多种恶意行为,相比现有对抗提示技术展现出显著优势。

原文摘要 · Abstract (English)

As large language models (LLMs) grow more capable, concerns about their safe deployment have also grown. Although alignment mechanisms have been introduced to deter misuse, they remain vulnerable to carefully designed adversarial prompts. In this work, we present a scalable attack strategy: intent-hiding adversarial prompting, which conceals malicious intent through the composition of skills. We develop a game-theoretic framework to model the interaction between such attacks and defense systems that apply both prompt and response filtering. Our analysis identifies equilibrium points and reveals structural advantages for the attacker. To counter these threats, we propose and analyze a defense mechanism tailored to intent-hiding attacks. Empirically, we validate the attack's effectiveness on multiple real-world LLMs across a range of malicious behaviors, demonstrating clear advantages over existing adversarial prompting techniques.

大模型安全对抗攻击博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。