arXiv:2509.05865cs.LGstat.ML2025-09

研究模型遗忘中的数据伪造问题,发现伪造样本极难生成。

The Measure of Deception: An Analysis of Data Forging in Machine Unlearning

  • 定义ε-伪造集:梯度接近目标的样本集合,用于分析伪造可能性。
  • 在高维数据下,伪造集测度随ε按幂律衰减,实际伪造极难。
  • 适用于线性模型与单层神经网络,为验证遗忘有效性提供理论依据。

受隐私法规和有害数据缓解需求推动,机器遗忘旨在修改训练好的模型,使其有效‘忘记’指定数据。验证遗忘的关键挑战在于‘伪造’——即恶意构造模仿目标点梯度的数据,制造已遗忘的假象。为此,本文提出ε-伪造集的概念:梯度在容忍度ε内逼近目标梯度的数据点集合,并建立其分析框架。对于线性回归和单层神经网络,证明该集合的Lebesgue测度很小,其大小与ε或ε^d成正比;在一般条件下,测度以ε^{(d−r)/2}速率衰减,其中d为数据维度,r<d为对应小奇异值右奇异向量张成空间的维数。对批量SGD及几乎处处光滑损失函数的扩展也得到相同渐近缩放。进一步,在非退化数据分布下,给出概率上界,表明随机采样伪造点的可能性可忽略不计。这些结果表明,对抗性伪造在本质上受限,虚假遗忘声明原则上可被检测。

原文摘要 · Abstract (English)

Motivated by privacy regulations and the need to mitigate the effects of harmful data, machine unlearning seeks to modify trained models so that they effectively ``forget'' designated data. A key challenge in verifying unlearning is \emph{forging} -- adversarially crafting data that mimics the gradient of a target point, thereby creating the appearance of unlearning without actually removing information. To capture this phenomenon, we consider the collection of data points whose gradients approximate a target gradient within tolerance $ε$ -- which we call an $ε$-forging set -- and develop a framework for its analysis. For linear regression and one-layer neural networks, we show that the Lebesgue measure of this set is small. It scales on the order of $ε$, and when $ε$ is small enough, $ε^d$. More generally, under mild regularity assumptions, we prove that the forging set measure decays as $ε^{(d-r)/2}$, where $d$ is the data dimension and $r<d$ is the dimension of vector space of right singular vectors corresponding to ``small'' singular values of a variation matrix defined by the model gradients. Extensions to batch SGD and almost-everywhere smooth loss functions yield the same asymptotic scaling. In addition, we establish probability bounds showing that, under non-degenerate data distributions, the likelihood of randomly sampling a forging point is vanishingly small. These results provide evidence that adversarial forging is fundamentally limited and that false unlearning claims can, in principle, be detected.

机器遗忘数据伪造模型安全理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。