用博弈论设计攻击,让大模型主动违规,成功率超95%。
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- 将攻击建模为可停止的随机博弈,用量化响应重写模型输出。
- 在囚徒困境等场景中使安全偏好失效,实现超过95%的越狱成功率。
- 适用于多语言、多模型,且能绕过关键词检测,适合安全研究者参考。
随着大语言模型普及,非专家用户可能带来风险,催生了大量越狱攻击研究。然而,现有黑盒越狱攻击多依赖手工设计或狭窄搜索空间,难以扩展。本文提出博弈论攻击(GTA),一种可扩展的黑盒越狱框架。我们将攻击者与安全对齐大模型的交互形式化为有限时域、可提前终止的随机序贯博弈,并通过量化响应重参数化模型的随机输出。基于此,我们提出‘模板-安全翻转’行为假设:通过博弈场景重构模型有效目标,原本的安全偏好会转变为最大化情景收益,从而在特定情境下弱化安全约束。我们在经典的囚徒困境变体等经典博弈中验证该机制,并引入攻击者代理,自适应升级压力以提升攻击成功率(ASR)。跨多个协议和数据集的实验表明,GTA在Deepseek-R1等模型上实现超95%的ASR,同时保持高效。消融实验验证了各组件的有效性与泛化能力。此外,场景扩展研究证明其可扩展性。GTA在其他博弈场景及单次生成的变体中也表现良好,且结合有害词检测代理后,在提示防护模型下仍维持高ASR。超越基准测试,GTA成功攻破真实世界的大模型应用,并对HuggingFace热门模型进行了长期安全监测。
原文摘要 · Abstract (English)
As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-crafted heuristics or narrow search spaces, which limit scalability. Compared with prior attacks, we propose Game-Theory Attack (GTA), an scalable black-box jailbreak framework. Concretely, we formalize the attacker's interaction against safety-aligned LLMs as a finite-horizon, early-stoppable sequential stochastic game, and reparameterize the LLM's randomized outputs via quantal response. Building on this, we introduce a behavioral conjecture "template-over-safety flip": by reshaping the LLM's effective objective through game-theoretic scenarios, the originally safety preference may become maximizing scenario payoffs within the template, which weakens safety constraints in specific contexts. We validate this mechanism with classical game such as the disclosure variant of the Prisoner's Dilemma, and we further introduce an Attacker Agent that adaptively escalates pressure to increase the ASR. Experiments across multiple protocols and datasets show that GTA achieves over 95% ASR on LLMs such as Deepseek-R1, while maintaining efficiency. Ablations over components, decoding, multilingual settings, and the Agent's core model confirm effectiveness and generalization. Moreover, scenario scaling studies further establish scalability. GTA also attains high ASR on other game-theoretic scenarios, and one-shot LLM-generated variants that keep the model mechanism fixed while varying background achieve comparable ASR. Paired with a Harmful-Words Detection Agent that performs word-level insertions, GTA maintains high ASR while lowering detection under prompt-guard models. Beyond benchmarks, GTA jailbreaks real-world LLM applications and reports a longitudinal safety monitoring of popular HuggingFace LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。