arXiv:2605.20654cs.LGcs.AI2026-05

让大模型在生成过程中自我反思,有效防御复杂越狱攻击。

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

论文配图:REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
图 1 · 摘自论文原文
  • 通过两阶段框架,在生成过程中内化自我反思能力。
  • 对复杂间接攻击的防御成功率超90%,且任务性能提升5.85%。
  • 适合关注模型安全与鲁棒性的研究者与开发者使用。

尽管大语言模型展现出强大能力,仍易受复杂多步越狱攻击影响,此类攻击通过利用内部生成过程绕过传统的表面安全对齐。为应对这一问题,我们提出Reflector,一种原理清晰的两阶段框架,将自我反思内化至生成轨迹中。首先,借助教师引导生成高质量反思数据,用于监督微调(SFT),建立结构化的反思模式;随后,采用以结果为导向、奖励有效性监督的强化学习(RL),培养出强健且自主的自我反思能力。实证结果表明,Reflector在对抗复杂间接攻击时防御成功率(DSR)超过90%,并在多种威胁场景下具有良好泛化性。值得注意的是,该框架同时提升了任务特定与通用能力,在GSM8K上取得5.85%的性能增益,并改善了知识密集型基准的表现。通过内化轨迹级安全机制,Reflector克服了表面对齐的根本局限,且无显著计算开销,为构建安全可靠的大型语言模型提供了高效可扩展的解决方案。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process. To address these vulnerabilities, we propose Reflector, a principled two-stage framework that internalizes self-reflection within the generation trajectory. Reflector first leverages teacher-guided generation to produce high-quality reflection data for supervised fine-tuning (SFT), establishing structured reflection patterns. It subsequently uses Reinforcement Learning (RL) with outcome-driven and reward-validity supervision to instill robust, autonomous self-reflection capabilities. Empirical results show that Reflector achieves Defense Success Rates (DSR) exceeding 90% against complex indirect attacks while generalizing robustly across diverse threat scenarios. Notably, the framework enhances both task-specific and general utility, yielding a 5.85% gain on GSM8K alongside improved performance on knowledge-intensive benchmarks. By internalizing trajectory-level safety, Reflector overcomes the fundamental limitations of surface alignment without significant computational overhead, offering an efficient and scalable solution for the development of safe and capable LLMs.

模型安全自我反思越狱防御强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。