揭示数据投毒攻击的隐蔽机制,提出防御与安全的不可兼得权衡。
Safety-Efficacy Trade Off: Robustness against Data-Poisoning
- 通过输入空间曲率分析,发现污染样本引发可检测的梯度特征。
- 在非线性核下,攻击成功率高但曲率趋零,导致无法被谱方法检测。
- 梯度正则化可抑制投毒,但会降低模型拟合能力,带来安全与性能权衡。
后门和数据投毒攻击能在逃避现有基于谱和优化的防御的同时实现高成功率。我们证明这种行为并非偶然,而是源于输入空间中的基本几何机制。利用核岭回归作为宽神经网络的精确模型,我们证明聚类的脏标签投毒会在输入海森矩阵中诱导出一个秩一尖峰,其幅度随攻击效能呈二次增长。关键的是,在非线性核下,我们识别出一种近似克隆态:此时攻击效能仍保持为常数阶,而诱导的输入曲率趋于零,使攻击在理论上无法通过谱手段检测。进一步表明,输入梯度正则化在梯度流下会压缩对齐投毒的费舍尔与海森特征模式,从而显式地、不可避免地造成安全与效能的权衡,降低数据拟合能力。对于指数核,该防御可精确解释为一种各向异性的高通滤波器,增加有效长度尺度并抑制近似克隆投毒。在线性模型和深度卷积网络上对MNIST、CIFAR-10和CIFAR-100的大量实验验证了理论,显示出攻击成功率与谱可见性之间的系统性滞后,并表明正则化与数据增强共同抑制投毒。结果首次完整刻画了投毒的可检测性、防御机制与输入空间曲率之间的关系。
原文摘要 · Abstract (English)
Backdoor and data poisoning attacks can achieve high attack success while evading existing spectral and optimisation based defences. We show that this behaviour is not incidental, but arises from a fundamental geometric mechanism in input space. Using kernel ridge regression as an exact model of wide neural networks, we prove that clustered dirty label poisons induce a rank one spike in the input Hessian whose magnitude scales quadratically with attack efficacy. Crucially, for nonlinear kernels we identify a near clone regime in which poison efficacy remains order one while the induced input curvature vanishes, making the attack provably spectrally undetectable. We further show that input gradient regularisation contracts poison aligned Fisher and Hessian eigenmodes under gradient flow, yielding an explicit and unavoidable safety efficacy trade off by reducing data fitting capacity. For exponential kernels, this defence admits a precise interpretation as an anisotropic high pass filter that increases the effective length scale and suppresses near clone poisons. Extensive experiments on linear models and deep convolutional networks across MNIST and CIFAR 10 and CIFAR 100 validate the theory, demonstrating consistent lags between attack success and spectral visibility, and showing that regularisation and data augmentation jointly suppress poisoning. Our results establish when backdoors are inherently invisible, and provide the first end to end characterisation of poisoning, detectability, and defence through input space curvature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。