arXiv:2602.02986cs.LGstat.ML2026-02

从线性稳定性角度揭示模型为何难删记忆,解释了数据对齐与遗忘难易的关系。

Why Some Models Resist Unlearning: A Linear Stability Perspective

  • 用渐近线性稳定性分析训练动态与数据几何的交互关系
  • 发现数据相干性决定遗忘能否成功,低信噪比更易遗忘
  • 为梯度类遗忘方法提供理论边界,适合研究隐私与模型可解释性者

机器遗忘是指在不重新训练的前提下擦除特定训练样本的影响,对隐私保护、合规性和效率至关重要。然而,当前大多数遗忘研究仍基于经验,缺乏对遗忘何时有效及其原因的理论理解。本文通过渐近线性稳定性视角,刻画优化过程与数据几何之间的相互作用。核心概念是数据相干性——即最优解附近损失曲面方向的跨样本对齐程度。我们沿保留集、遗忘集及其之间三个维度分解相干性,并推导出收敛与发散的精确稳定性阈值。进一步在信号加噪声模型下的两层ReLU CNN中分析,发现更强的记忆化反而使遗忘更容易:当信噪比(SNR)较低时,跨样本对齐较弱,相干性降低,遗忘更易实现;反之,高SNR下高度对齐的模型抵抗遗忘。实证验证显示,海森测试与CNN热图紧密契合预测边界,映射出基于批次、混合策略和数据/模型对齐的梯度遗忘的稳定性前沿。分析基于随机矩阵理论工具,首次提供了关于记忆化、相干性与遗忘之间权衡的原理性解释。

原文摘要 · Abstract (English)

Machine unlearning, the ability to erase the effect of specific training samples without retraining from scratch, is critical for privacy, regulation, and efficiency. However, most progress in unlearning has been empirical, with little theoretical understanding of when and why unlearning works. We tackle this gap by framing unlearning through the lens of asymptotic linear stability to capture the interaction between optimization dynamics and data geometry. The key quantity in our analysis is data coherence which is the cross sample alignment of loss surface directions near the optimum. We decompose coherence along three axes: within the retain set, within the forget set, and between them, and prove tight stability thresholds that separate convergence from divergence. To further link data properties to forgettability, we study a two layer ReLU CNN under a signal plus noise model and show that stronger memorization makes forgetting easier: when the signal to noise ratio (SNR) is lower, cross sample alignment is weaker, reducing coherence and making unlearning easier; conversely, high SNR, highly aligned models resist unlearning. For empirical verification, we show that Hessian tests and CNN heatmaps align closely with the predicted boundary, mapping the stability frontier of gradient based unlearning as a function of batching, mixing, and data/model alignment. Our analysis is grounded in random matrix theory tools and provides the first principled account of the trade offs between memorization, coherence, and unlearning.

机器遗忘线性稳定性数据相干性信噪比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。