arXiv:2605.02109cs.LGcs.CR2026-05

通过数学证明增强对抗噪声,实现无需训练的高效检测。

Detecting Adversarial Data via Provable Adversarial Noise Amplification

论文配图:Detecting Adversarial Data via Provable Adversarial Noise Amplification
图 1 · 摘自论文原文
  • 基于谱损失与架构设计,强制网络放大对抗噪声
  • 在多个攻击下检测准确率超95%,对抗自适应攻击也有效
  • 轻量级推理检测,适合部署于资源受限场景

深度神经网络中对抗噪声在各层间非均匀且持续放大的现象,虽被广泛用于检测对抗样本并提升鲁棒性,但缺乏严格的数学依据。本文深入研究该现象,提出可证明的对抗噪声放大定理,明确在特定条件下噪声放大可被数学保证。基于此理论,我们设计了一种新型训练方法:采用定制谱损失函数和特定网络结构以增强放大信号。最后,提出一种全新的轻量级检测机制,仅在推理阶段运行,利用强化后的放大信号进行判断。实验验证了该方法对当前主流攻击及专为规避检测设计的自适应攻击均具有效性,表明增强的噪声放大可作为可靠、稳健的对抗防御信号。

原文摘要 · Abstract (English)

The nonuniform and growing impact of adversarial noise across the layers of deep neural networks has been used in the literature, without a formal mathematical justification, to detect adversarial inputs and improve robustness. In this work, we study this phenomenon in detail and present a formal adversarial noise amplification theorem. We specify a set of sufficient conditions under which the adversarial noise amplification is mathematically guaranteed. Based on theoretical observations, we propose a novel training methodology with a custom spectral loss function and a specific architectural design to enhance the amplification signal for detecting adversarial data. Finally, we introduce a new, lightweight detection mechanism that leverages the enhanced amplification signal and operates entirely at inference time. To validate our approach, we demonstrate the detector's efficacy against both state-of-the-art attacks and a purpose-built adaptive attack, confirming that enhanced amplification can serve as a robust and reliable signal for adversarial defense.

对抗检测噪声放大推理检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。