通过隐空间对齐提升大模型推理安全性,有效防御越狱攻击。
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
- 利用隐状态空间优化,让模型生成更安全的推理过程。
- 在多个基准上使推理安全提升79.0%,最终响应安全提升87.7%。
- 适合关注大模型安全对齐与对抗攻击防护的研究者。
我们提出CRAFT,一种基于红队测试的对齐框架,利用模型推理能力与隐表示提升对抗越狱攻击的鲁棒性。不同于以往主要在输出层面进行防御的方法,CRAFT通过显式优化隐状态空间中的目标,使大推理模型生成具有安全意识的推理轨迹。方法上,融合对比表征学习与强化学习,分离安全与非安全推理路径,构建支持鲁棒推理级对齐的潜在空间几何。理论上,将隐-文一致性引入GRPO可排除表面对齐策略作为局部最优解。实验中,我们在Qwen3-4B-Thinking与R1-Distill-Llama-8B两个强推理模型上评估CRAFT,在多个安全基准上持续优于IPO和SafeKey等先进防御方法。显著地,相比基线模型,推理安全平均提升79.0%,最终响应安全提升87.7%,验证了隐空间推理对齐的有效性。
原文摘要 · Abstract (English)
We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness against jailbreak attacks. Unlike prior defenses that operate primarily at the output level, CRAFT aligns large reasoning models to generate safety-aware reasoning traces by explicitly optimizing objectives defined over the hidden state space. Methodologically, CRAFT integrates contrastive representation learning with reinforcement learning to separate safe and unsafe reasoning trajectories, yielding a latent-space geometry that supports robust, reasoning-level safety alignment. Theoretically, we show that incorporating latent-textual consistency into GRPO eliminates superficially aligned policies by ruling them out as local optima. Empirically, we evaluate CRAFT on multiple safety benchmarks using two strong reasoning models, Qwen3-4B-Thinking and R1-Distill-Llama-8B, where it consistently outperforms state-of-the-art defenses such as IPO and SafeKey. Notably, CRAFT delivers an average 79.0% improvement in reasoning safety and 87.7% improvement in final-response safety over the base models, demonstrating the effectiveness of hidden-space reasoning alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。