arXiv:2606.20508cs.AIcs.LG2026-06

研究安全对齐大模型如何理解好坏示范的混合效果

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

论文配图:What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
图 1 · 摘自论文原文
  • 混合使用良性与有害示范,测试模型对合规性的学习机制
  • 良性示范可能降低或反而提升有害行为,取决于模型类型
  • 偏好优化阶段是防止良性示范助长有害行为的关键

已有研究显示,上下文示范可导致语言模型被越狱,但模型如何理解不同类型的合规示范仍不明确。本文通过混合良性合规示范(非有害请求,有益回复)与有害合规示范(有害请求,有益回复),测试三种关于示范组合如何驱动有害合规的假设。在四个模型上发现,良性与有害示范不可互换:良性示范可能降低或增加有害合规,具体取决于模型。进一步表明,偏好优化是防止良性示范加剧有害行为的关键训练阶段;示范顺序存在显著近期效应;模型在拒绝时对上下文学习的响应也不同:部分模型即使拒绝仍会模仿示范格式,另一些则完全忽略上下文信号。综上,该研究从展示示范越狱有效,转向揭示其内在机制:模型从合规示范中提取的信息取决于示范内容、顺序及训练方法。

原文摘要 · Abstract (English)

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrations are not interchangeable: benign demonstrations can either reduce or increase harmful compliance depending on the model. We further show that preference optimization is the critical training stage that prevents benign demonstrations from increasing harmful compliance, that demonstration ordering exhibits strong recency bias, and that models differ in how refusal interacts with in-context learning: some adopt demonstrated formatting even when refusing, while others override all in-context signals upon refusal. Taken together, this work moves beyond showing that demonstration-based jailbreaking works to characterizing how it works: what models extract from compliance demonstrations depends on demonstration content, ordering, and training methodology.

大模型安全示范学习合规性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。