arXiv:2606.30263cs.CRcs.AI2026-06

隐藏有害指令的伪装样本能绕过现有防御,新方法有效应对。

Defending Against Harmful Supervision Hidden in Benign Samples

论文配图:Defending Against Harmful Supervision Hidden in Benign Samples
图 1 · 摘自论文原文
  • 将有害问答嵌入正常数据中制造隐蔽攻击
  • 主流防护机制在样本级检测失效
  • 通过分词级正则化提升微调安全性,适合安全敏感场景

现有防御在有害内容显式混入微调数据时有效,但攻击者可将有害监督信息隐藏于良性任务样本中。我们提出嵌入式攻击(Embedded Attack),验证主流护栏在样本层面常无法检测此类隐藏威胁。为应对该问题,我们提出双参考监督微调(DR-SFT),借鉴DPO的对比目标设计,通过分词级正则化适配SFT,显著缓解粗粒度数据过滤之外的有害微调风险。

原文摘要 · Abstract (English)

Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign tasks. We propose Embedded Attack, where harmful QA pairs are embedded within benign training samples, and show that representative guardrails often fail to detect them at the example level. To address this, we propose Dual-Reference SFT (DR-SFT), which adapts DPO-style contrastive objective design to SFT through token-level regularization, mitigating harmful fine-tuning beyond coarse data filtering.

模型安全微调防御隐蔽攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。