arXiv:2509.24393cs.AIcs.CL2025-09被引 10

通过干预式优化提升大模型推理过程的安全性,防止有害内容泄露。

Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention

  • 用关键安全触发点替换合规步骤,增强推理过程的可引导性。
  • 在对抗性测试中,有害内容减少超30%,且保持原有推理能力。
  • 适合关注模型安全、防范恶意攻击的研究者与应用开发者。

尽管大型推理模型(LRMs)在解决复杂问题上取得进展,其思维链(CoT)推理中仍常包含有害内容,即使最终输出看似安全。现有方法忽视了推理过程本身的安全性,导致可信度下降,可能被恶意用户利用。本文聚焦于对推理过程的安全对齐,提出干预式偏好优化(IPO)。研究发现:1)安全推理往往由少数关键安全触发步骤决定;2)合规提示与不安全延续强相关;3)纠正性干预能有效引导不安全路径回归安全轨迹。IPO通过替换合规步骤为安全触发,并构建强信号偏好对进行训练。在越狱和对抗性安全基准测试中,IPO显著提升推理与输出的整体安全性,相比SFT与强化学习基线,有害内容相对减少超30%,同时在多样化推理任务中保持优异表现。结果凸显显式推理对齐的重要性,为构建更安全的LRMs提供了可行路径。

原文摘要 · Abstract (English)

Although Large Reasoning Models (LRMs) have progressed in solving complex problems, their chain-of-thought (CoT) reasoning often contains harmful content that can persist even when the final responses appear safe. We show that this issue still remains in existing methods which overlook the unique significance of safe reasoning, undermining their trustworthiness and posing potential risks in applications if unsafe reasoning is accessible for and exploited by malicious users. We therefore shift our focus to aligning the safety of reasoning itself in this paper and explore process supervision as the solution. However, simply rewarding safe reasoning proves inadequate due to low rollout diversity and limited training signals. To tackle this challenge, we first delve into the characteristics of safe reasoning and uncover several critical insights that 1) safe reasoning is often consolidated by a few critical steps of safety triggers; 2) compliance cues strongly correlate with unsafe continuations; and 3) corrective interventions reliably steer unsafe trajectories towards safer traces. Motivated by these, we propose Intervened Preference Optimization (IPO), an alignment method that enforces safe reasoning by substituting compliance steps with safety triggers and constructing pairs for preference learning with strong signals. Experiments on jailbreak and adversarial safety benchmarks demonstrate that IPO remarkably improves overall safety regarding both reasoning and responses, outperforming SFT-based and RL-based baselines with a relative reduction of over 30% in harmfulness, while preserving excellent performance across diverse reasoning tasks. The results highlight the importance of explicit alignment for reasoning and provide a practical path to safer LRMs.

模型安全推理对齐对抗防御偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。