arXiv:2605.01034cs.CL2026-05

构建攻击与防御的理论博弈模型,揭示攻击者天然优势并推导出最优防御策略。

A Theoretical Game of Attacks via Compositional Skills

论文配图:A Theoretical Game of Attacks via Compositional Skills
图 1 · 摘自论文原文
  • 将攻击与防御建模为博弈,设计理论最优攻击策略。
  • 证明攻击方在均衡中具有天然优势,且新攻击方法性能更优。
  • 提出可证明最优的防御方案,适用于多种大模型和评测场景。

随着大语言模型能力不断提升,其安全部署问题日益突出。尽管已有诸多对齐策略试图限制有害行为,但精心设计的对抗性提示仍可绕过这些防御。本文提出一个攻击者与防御者之间的理论博弈框架,设计出理论上最优的响应攻击策略,并发现其与多种现有对抗提示方法密切相关。进一步分析表明,该博弈存在均衡,且攻击方具有固有优势。基于理论分析,我们推导出一个可证明最优的防御策略。实验上,我们评估了该理论最优攻击的实际实现,在不同大模型和基准测试中均展现出优于现有对抗提示方法的表现。

原文摘要 · Abstract (English)

As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we introduce a theoretical framework that formalizes a game between an attacker and a defender. Within this framework, we design a theoretical best-response attack strategy and show that it is closely related to many existing adversarial prompting methods. We further analyze the resulting game, characterize its equilibria, and reveal inherent advantages for the attacker. Drawing on our theoretical analysis, we also derive a provably optimal defense strategy. Empirically, we evaluate a practical instantiation of the theoretically optimal attack and observe stronger performance relative to existing adversarial prompting approaches in diverse settings encompassing different LLMs and benchmarks.

对抗攻击博弈论大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。