arXiv:2508.12897cs.AIcs.CR2025-08

用具体提示激发模型推理出有害内容,再通过原则引导修复,提升大模型安全性和推理能力。

RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models

  • 通过具体化恶意提示触发模型逐步推理出有害内容,暴露安全漏洞。
  • 构建包含3989条样本的PGA数据集,使模型防御成功率提升29.5%。
  • 既增强安全性又不削弱推理能力,适合需要高可靠性的推理模型应用。

大型推理模型(LRMs)存在一种独特安全漏洞:其内部推理链可能生成有害内容,而最终输出看似无害。为解决这一被忽视的风险,我们提出新型攻击范式——通过具体化触发的推理激活越狱(RAJ),证明将恶意提示细化可诱导模型产生绕过安全协议的逻辑推理链。为进一步系统缓解该漏洞,我们开发了一种可扩展的安全对齐数据集构建框架:先利用RAJ攻击从LRMs中提取高风险有害推理链,再通过定制的原理引导对齐(PGA)机制将其转化为安全、建设性且具教育意义的回应。随后,我们构建了PGA数据集,共包含3,989个样本。大量实验表明,使用PGA数据集微调后,模型在多个越狱基准测试中的防御成功率最高提升29.5%。关键的是,该方法不仅有效抵御复杂推理型攻击,还保持甚至增强了模型的通用推理能力。本工作为推理密集型AI系统的安全对齐提供了可扩展、高效的新路径,解决了安全性与功能性之间的核心矛盾。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) face a distinct safety vulnerability: their internal reasoning chains may generate harmful content even when the final output appears benign. To address this overlooked risk, we first propose a novel attack paradigm, Reasoning-Activated Jailbreak (RAJ) via Concretization, which demonstrates that refining malicious prompts to be more specific can trigger step-by-step logical reasoning that overrides the model's safety protocols. To systematically mitigate this vulnerability, we further develop a scalable framework for constructing high-quality safety alignment datasets. This framework first leverages the RAJ attack to elicit challenging harmful reasoning chains from LRMs, then transforms these high-risk traces into safe, constructive, and educational responses through a tailored Principle-Guided Alignment (PGA) mechanism. Then, we introduce the PGA dataset, a verified alignment dataset containing 3,989 samples using our proposed method. Extensive experiments show that fine-tuning LRMs with PGA dataset significantly enhances model safety, achieving up to a 29.5% improvement in defense success rates across multiple jailbreak benchmarks. Critically, our approach not only defends against sophisticated reasoning-based attacks but also preserves, even enhances, the model's general reasoning capabilities. This work provides a scalable and effective pathway for safety alignment in reasoning-intensive AI systems, addressing the core trade-off between safety and functional performance.

模型安全推理对齐越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。