用博弈论建模提示攻击与大模型防御,提升安全防护能力
Toward a Dynamic Stackelberg Game-Theoretic Framework for Agentic AI Defense Against LLM Jailbreaking
- 将攻击者与大模型互动建模为带RRT搜索的动态博弈
- 识别出攻击者无法获利的防御平衡点,实现有效拦截
- 适合研究大模型安全、对抗攻击与防御机制的学者
本文提出一种博弈论框架,将提示工程师与大语言模型(LLMs)的交互建模为两方的扩展式博弈,并结合快速探索随机树(RRT)在提示空间中进行搜索。攻击者逐步采样、扩展并测试提示,而模型则选择接受、拒绝或引导,导致最终结果为安全交互、被阻断或越狱。将RRT探索嵌入扩展式博弈中,同时捕捉了越狱策略的发现过程和模型的战略响应。此外,我们证明防御者行为可通过局部斯塔克尔伯格均衡解释,说明攻击者无法再获得有利的提示偏离,为紫队代理防御的有效性提供理论依据。由此生成的游戏树为评估、解释和强化大模型防护机制提供了原则性基础。
原文摘要 · Abstract (English)
This paper proposes a game theoretic framework that models the interaction between prompt engineers and large language models (LLMs) as a two player extensive form game coupled with a Rapidly exploring Random Trees (RRT) search over prompt space. The attacker incrementally samples, extends, and tests prompts, while the LLM chooses to accept, reject, or redirect, leading to terminal outcomes of Safe Interaction, Blocked, or Jailbreak. Embedding RRT exploration inside the extensive form game captures both the discovery phase of jailbreak strategies and the strategic responses of the model. Furthermore, we show that the defender behavior can be interpreted through a local Stackelberg equilibrium condition, which explains when the attacker can no longer obtain profitable prompt deviations and provides a theoretical lens for understanding the effectiveness of our Purple Agent defense. The resulting game tree thus offers a principled foundation for evaluating, interpreting, and hardening LLM guardrails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。