arXiv:2602.04899cs.CRcs.AI2026-02被引 4

一种可绕过所有数据级防御的数据投毒攻击,能长期潜伏并触发特定行为。

Phantom Transfer: Data Poisoning can Survive Data-Level Defences

  • 利用隐性学习机制设计新型投毒,即使知晓投毒方式也无法过滤。
  • 在11种数据级防御下仍有效,包括模型重写样本的防护。
  • 适用于多种模型和任务,适合研究防御漏洞与后训练审计者参考。

我们提出一种数据投毒攻击——幻影迁移(Phantom Transfer),其特性在于:即便确切知道毒物如何被注入原本无害的数据集,也无法将其过滤掉。通过使隐性学习适应真实场景,该攻击在任意生成数据的模型、任意训练模型及任意目标下均有效。此外,该攻击在11种已测试的数据级防御中存活,其中包括每条样本均由另一模型改写的情况。我们分析了攻击最有效的条件,并证明其可用于植入密码触发的行为,同时突破防御。简言之,本工作提供了一个存在性证明:最高权限的防御仍可能无法阻止复杂的数据投毒攻击。我们建议未来防御应结合白盒方法与训练后模型审计。

原文摘要 · Abstract (English)

We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produced the data, which model is trained on the data or what the attack target is. Furthermore, the attack survives 11 tested data-level defences, including one where every sample is paraphrased by another model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. In short, we provide an existence proof that maximum-affordance defences can fail to stop sophisticated data poisoning attacks. We suggest that future defences should be supplemented with white-box methods and post-training model audits.

数据投毒安全防御模型鲁棒性白盒审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。