arXiv:2510.00938cs.LG2025-10被引 10

让大模型学会纠正错误推理,提升安全性和抗攻击能力。

Large Reasoning Models Learn Better Alignment from Flawed Thinking

  • 用反向思维提示+强化学习,教会模型自我纠错
  • 安全防护力提升,越狱攻击成功率下降40%以上
  • 适合需要高可靠性的AI助手、内容审核等场景

大型推理模型(LRMs)通过生成结构化思维链(CoT)来推理解答问题,但面对错误前提时仍缺乏批判性反思能力,易受偏见影响。本文提出RECAP(基于反向对齐预填充的鲁棒安全对齐),一种无需额外训练成本或修改的强化学习方法,通过混合合成的反向对齐思维链与标准提示进行后训练,使模型能够主动识别并绕过错误推理路径,转向安全且有帮助的回答。实验表明,RECAP显著提升了安全性与越狱攻击抵御能力,减少过度拒绝现象,同时保持核心推理能力不变,且不增加推理阶段的词元开销。深入分析显示,经过RECAP训练的模型更频繁进行自我反思,在持续对抗攻击下仍能维持安全表现。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flawed premise is injected into their thought process. We propose RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses. RECAP trains on a mixture of synthetically generated counter-aligned CoT prefills and standard prompts, requires no additional training cost or modifications beyond vanilla reinforcement learning from human feedback (RLHF), and substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability -- all while maintaining inference token budget. Extensive analysis shows that RECAP-trained models engage in self-reflection more frequently and remain robust under adaptive attacks, preserving safety even after repeated attempts to override their reasoning.

大模型安全对齐推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。