arXiv:2602.04224cs.LGcs.AI2026-02被引 1

提出风险感知的偏好优化框架,让大模型更安全地应对复杂越狱攻击。

RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning

  • 基于风险感知的偏好优化,动态识别推理中的安全风险。
  • 在多种越狱攻击下,安全推理泛化能力提升,且不影响通用性能。
  • 适合关注大模型安全对齐的研究者与开发者使用。

大型推理模型(LRM)虽在思维链(CoT)推理上取得显著进展,但仍面临与基础语言模型相似的安全问题。现有方法虽能引导模型拒绝有害提示,但对多样且复杂的越狱攻击泛化能力不足。本文指出,安全推理的泛化失败源于其过程的不充分性,尤其在复杂攻击提示面前表现不佳。我们通过理论与实证证明,需更充分的安全推理过程以抵御高级攻击。为此,提出风险感知偏好优化(RAPO)框架,使模型能在推理中自适应识别并以适当粒度应对安全风险。大量实验表明,RAPO能有效提升多种LRM在多样化攻击提示下的安全推理泛化能力,同时保持通用性能,为大模型安全对齐提供稳健技术。代码已开源:https://github.com/weizeming/RAPO。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have achieved tremendous success with their chain-of-thought (CoT) reasoning, yet also face safety issues similar to those of basic language models. In particular, while algorithms are designed to guide them to deliberately refuse harmful prompts with safe reasoning, this process often fails to generalize against diverse and complex jailbreak attacks. In this work, we attribute these failures to the generalization of the safe reasoning process, particularly their insufficiency against complex attack prompts. We provide both theoretical and empirical evidence to show the necessity of a more sufficient safe reasoning process to defend against advanced attack prompts. Building on this insight, we propose a Risk-Aware Preference Optimization (RAPO) framework that enables LRM to adaptively identify and address the safety risks with appropriate granularity in its thinking content. Extensive experiments demonstrate that RAPO successfully generalizes multiple LRMs' safe reasoning adaptively across diverse attack prompts whilst preserving general utility, contributing a robust alignment technique for LRM safety. Our code is available at https://github.com/weizeming/RAPO.

大模型安全推理对齐越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。