提出线性模型数据投毒的定量分析框架,揭示攻击强度与模型参数的关系。
A Linear Approach to Data Poisoning
- 基于高维随机矩阵理论,推导出投毒后模型得分的闭式极限表达式。
- 发现当过参数化程度趋近1时出现插值阈值,权重会与投毒方向对齐。
- 适用于理解线性模型中的后门攻击,适合研究模型安全与鲁棒性的人参考。
后门攻击和数据投毒攻击可在极小训练扰动下改变预测结果,但缺乏关于投毒强度、过参数化与正则化之间关系的清晰理论。本文在高维情形下分析带有未惩罚截距项的岭回归最小二乘法(p,n→∞,p/n→c)。通过将目标投毒建模为将某一类θ比例样本沿方向v移动并重新标记,结合再生核技术与随机矩阵理论的确定性等价物,推导出中毒后得分的闭式极限表达式,其显式依赖于模型参数。公式揭示了缩放规律,恢复了无正则化极限下c→1时的插值阈值,并表明权重与投毒方向对齐。合成实验在参数扫描中与理论高度一致,且在MNIST后门测试中显示出定性上一致的趋势。结果为量化线性模型中的投毒行为提供了可解析的框架。
原文摘要 · Abstract (English)
Backdoor and data-poisoning attacks can flip predictions with tiny training corruptions, yet a sharp theory linking poisoning strength, overparameterization, and regularization is lacking. We analyze ridge least squares with an unpenalized intercept in the high-dimensional regime \(p,n\to\infty\), \(p/n\to c\). Targeted poisoning is modelled by shifting a \(θ\)-fraction of one class by a direction \(\mathbf{v}\) and relabelling. Using resolvent techniques and deterministic equivalents from random matrix theory, we derive closed-form limits for the poisoned score explicit in the model parameters. The formulas yield scaling laws, recover the interpolation threshold as \(c\to1\) in the ridgeless limit, and show that the weights align with the poisoning direction. Synthetic experiments match theory across sweeps of the parameters and MNIST backdoor tests show qualitatively consistent trends. The results provide a tractable framework for quantifying poisoning in linear models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。