arXiv:2607.09532cs.LGcs.CR2026-07

攻击者可植入无法检测的后门,让模型生成特定对抗样本。

Statistically Undetectable Backdoors in Deep Neural Networks

论文配图:Statistically Undetectable Backdoors in Deep Neural Networks
图 1 · 摘自论文原文
  • 通过构造统计不可检测的后门,使模型在白盒下与正常模型难以区分。
  • 后门可对任意输入生成基于不变性的对抗样本,但无后门时无法在多项式时间内实现。
  • 揭示了训练者与使用者之间的根本能力不对称,适合安全研究者关注。

我们展示了攻击者如何在一大类深度前馈神经网络中植入后门。这些后门在白盒设置下是统计不可检测的,即即便已知模型完整结构(如所有权重),后门模型与正常训练模型之间的总变差距离也很小。该后门可为每个输入提供基于不变性的对抗样本,将相距较远的输入映射到异常接近的输出。然而,在没有后门的情况下,根据标准密码学假设,在多项式时间内证明无法生成此类对抗样本。理论分析与初步实证结果表明,模型训练者与使用者之间存在根本性能力不对称。

原文摘要 · Abstract (English)

We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.

后门攻击对抗样本模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。