通过注意力引导强化学习,提升对大推理模型的越狱攻击成功率。
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models

- 用强化学习优化越狱攻击,奖励函数结合注意力分布信号。
- 在三个基准上对五种模型测试,攻击成功率显著高于现有方法。
- 适合研究模型安全与对抗攻击的人员阅读。
大型推理模型(LRMs)通过生成结构化、分步推理内容,在解决复杂问题方面展现出强大能力。然而,暴露模型内部推理过程会引入新的安全风险;近期研究表明,相较于标准大语言模型,LRMs 更易遭受越狱攻击。本文研究了针对 LRMs 的越狱攻击,发现攻击成功率(ASR)与模型注意力模式密切相关:成功的越狱通常对输入提示中的有害标记分配较低注意力,而对推理内容中的相关标记分配更高注意力。受此启发,我们提出一种新型越狱方法,利用强化学习(RL)提升攻击效果,将注意力信号显式融入奖励函数设计。此外,引入多样化的说服策略以丰富 RL 动作空间,持续提升攻击成功率。在三个基准上对五种开源与闭源 LRMs 的广泛实验表明,该方法在有效性、效率和迁移性方面均显著优于现有方法。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in solving complex problems by generating structured, step-by-step reasoning content. However, exposing a model's internal reasoning process introduces additional safety risks; for example, recent studies show that LRMs are more vulnerable to jailbreak attacks than standard LLMs. In this paper, we investigate jailbreak attacks on LRMs and reveal that the attack success rate (ASR) is closely correlated with LRMs' attention patterns. Specifically, successful jailbreaks tend to assign lower attention to harmful tokens in the input prompt, while allocating higher attention to those tokens in the reasoning content. Motivated by this finding, we propose a novel jailbreak method for LRMs that leverages reinforcement learning (RL) to enhance attack effectiveness, explicitly incorporating attention signals into the reward function design. In addition, we introduce diverse persuasion strategies to enrich the RL action space, which consistently improves the ASR. Extensive experiments on five open-source and closed-source LRMs across three benchmarks demonstrate that our method achieves substantially higher ASR, outperforming existing approaches in terms of effectiveness, efficiency, and transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。