arXiv:2604.02652cs.LGcs.AI2026-04中稿 · JSAI 2026

发现大模型对齐存在泛化缺陷,复合攻击可大幅突破安全防护

Generalization Limits of Reinforcement Learning Alignment

  • 设计复合攻击手段,融合多种已知防御的漏洞
  • 对gpt-oss-20b攻击成功率从14.3%提升至71.4%
  • 揭示对齐训练泛化能力弱于模型能力,适合安全评估研究者

大型语言模型的安全依赖于基于人类反馈的强化学习(RLHF)等对齐技术。然而,近期理论分析表明,基于强化学习的训练并未带来新能力,仅重新分配了已有能力的使用概率。本文针对OpenAI gpt-oss-20b提出“复合越狱”攻击,利用对齐机制的泛化失败,将多种单独可防御的攻击技术组合,以饱和指令层级维护过程。实验显示,单一方法攻击成功率(ASR)为14.3%,而组合攻击达到71.4%。结果为对齐训练泛化能力弱于模型能力提供了实证支持,强调需采用复合攻击场景进行多维度安全评估。

原文摘要 · Abstract (English)

The safety of large language models (LLMs) relies on alignment techniques such as reinforcement learning from human feedback (RLHF). However, recent theoretical analyses suggest that reinforcement learning-based training does not acquire new capabilities but merely redistributes the utilization probabilities of existing ones. In this study, we propose ``compound jailbreaks'' targeting OpenAI gpt-oss-20b, which exploit the generalization failures of alignment. This approach combines multiple attack techniques -- each individually defended against -- to saturate the instruction hierarchy maintenance process. Our evaluation shows that the attack success rate (ASR) increased from 14.3\% with individual methods to 71.4\% with the combined approach. These results provide empirical evidence for the hypothesis that safety training does not generalize as broadly as model capabilities, highlighting the need for multifaceted safety evaluations using compound attack scenarios.

强化学习模型安全越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。