arXiv:2605.14605cs.CRcs.AI2026-05被引 2

现有防御方法被新攻击破解,暴露了安全漏洞。

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries

论文配图:One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
图 1 · 摘自论文原文
  • 提出统一自适应攻击,突破15种防御机制
  • 实测所有防御在新攻击下均失效
  • 适合模型安全研究者与部署实践者参考

模型提供商日益开放权重或允许通过API微调基础模型。尽管这些模型发布前已对齐安全,但通过有害数据微调仍可移除其防护机制。近期防御试图使模型抵御恶意微调,但主要针对固定攻击评估,未考虑防御存在时的对抗行为。我们调查了15种近期防御,发现其共性弱点:仅掩盖或误导有害行为路径,未消除行为本身。进而提出一种统一自适应攻击,成功攻破所有防御机制。结果表明,当前方法无法提供真正鲁棒的安全性,仅能防范其设计所针对的攻击。我们希望该统一自适应对手能帮助未来研究者和从业者在部署前充分测试新防御。

原文摘要 · Abstract (English)

Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. Recent defenses aim to make models robust to such malicious fine-tuning, but they are largely evaluated only against fixed attacks that do not account for the defense. We show that these robustness claims are incomplete. Surveying 15 recent defenses, we identify several defense mechanisms and show that they share a single weakness: they obscure or misdirect the path to harmful behavior without removing the behavior itself. We then develop a unified adaptive attack that breaks defenses across all defense mechanisms. Our results show that current approaches do not provide robust security; they mainly stop the attacks they were designed against. We hope that our unified adaptive adversary for this domain will help future researchers and practitioners stress-test new defenses before deployment.

模型安全对抗攻击微调防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。