强触发器反噬:高维下更强的训练触发反而提升防御效果
When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks
- 在高维数据中,通过调节训练触发强度,发现攻击效果非线性变化
- 攻击成功率在特定强度达到峰值后下降,清洁测试准确率随强度上升
- 最危险的触发方向是数据协方差的最小特征向量,适合安全研究者参考
在高维情形下,后门投毒攻击表现出反直觉现象:更强的训练触发反而有利于防御。本文研究在比例极限($p/n \to κ$)下的正则化广义线性模型,针对高斯混合数据,调节训练触发强度 $α$ 而固定测试触发。发现三种现象:(i) 清洁测试准确率随 $α$ 增加而提高;(ii) 攻击成功率在有限 $α$ 处达峰值后下降;(iii) 最具破坏性的触发方向为数据协方差矩阵的最小特征向量。对平方损失,三者均给出闭式解析证明,并将 (i)(ii) 推广至一般凸 GLM 损失,基于高斯代理不动点系统。我们揭示了与 $κ$ 成正比的有限样本噪声底限是 (i) 的机制来源,经典 $n \gg p$ 分析无法捕捉。在 CIFAR-10 和高斯模拟数据上的实验与理论高度吻合;ResNet-18 实验表明该现象在非凸设置下依然存在。
原文摘要 · Abstract (English)
Backdoor poisoning attacks behave counter-intuitively in high dimensions: stronger training triggers can help the defender. We study regularised generalised linear models on Gaussian-mixture data in the proportional regime ($p/n \to κ$), varying the training trigger strength $α$ against a fixed test trigger. Three phenomena emerge: (i) clean test accuracy increases with $α$; (ii) attack success peaks at a finite $α$ and then declines; and (iii) the most damaging trigger direction is the minimum eigenvector of the data covariance. We prove all three results in closed form for the squared loss, and extend (i) and (ii) to general convex GLM losses via a Gaussian-proxy fixed-point system. We identify a finite-sample noise floor proportional to $κ$ as the mechanism behind (i), invisible to classical $n \gg p$ analysis. Experiments on CIFAR-10 and Gaussian surrogates match the theory closely; ResNet-18 experiments show the same phenomena beyond the convex setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。