高维机器遗忘中,单步牛顿法加高斯噪声即可同时保证隐私与准确。
Gaussian Certified Unlearning in High Dimensions: A Hypothesis Testing Approach
- 提出ε-高斯可认证性,适配高维场景的噪声添加机制。
- 证明单步牛顿法+校准高斯噪声足以实现隐私与精度。
- 挑战旧结论,指出其认证标准不兼容实际噪声机制。
机器遗忘旨在高效消除特定数据的影响同时保持模型泛化能力。低维情形($p \ll n$)已取得显著进展,但高维下因单样本损失函数 $f$ 的强凸性 $Ω(1)$ 和光滑性 $O(1)$ 标准假设难以同时成立,带来严重理论挑战。本文引入ε-高斯可认证性,一种在高维设定中自然且稳健的概念,能最优刻画广泛噪声添加机制。我们对广为使用的基于单步牛顿法的遗忘算法进行理论分析,结果表明:在所述高维设置下,仅需一次牛顿步配合经校准的高斯噪声,即可同时实现隐私保护与模型准确率。该结论与先前唯一研究高维遗忘的工作(Zou et al., 2025)形成鲜明对比——后者放宽标准优化假设,采用ε-可认证性,却得出至少需两步才能保障隐私与准确性的结论。我们推断此差异源于ε-可认证性本身的非最优性及其与噪声机制的不兼容性,而ε-高斯可认证性则能有效克服这一缺陷。
原文摘要 · Abstract (English)
Machine unlearning seeks to efficiently remove the influence of selected data while preserving generalization. Significant progress has been made in low dimensions $(p \ll n)$, but high dimensions pose serious theoretical challenges as standard optimization assumptions of $Ω(1)$ strong convexity and $O(1)$ smoothness of the per-example loss $f$ rarely hold simultaneously in proportional regimes $(p\sim n)$. In this work, we introduce $\varepsilon$-Gaussian certifiability, a canonical and robust notion well-suited to high-dimensional regimes, that optimally captures a broad class of noise adding mechanisms. Then we theoretically analyze the performance of a widely used unlearning algorithm based on one step of the Newton method in the high-dimensional setting described above. Our analysis shows that a single Newton step, followed by a well-calibrated Gaussian noise, is sufficient to achieve both privacy and accuracy in this setting. This result stands in sharp contrast to the only prior work that analyzes machine unlearning in high dimensions \citet{zou2025certified}, which relaxes some of the standard optimization assumptions for high-dimensional applicability, but operates under the notion of $\varepsilon$-certifiability. That work concludes %that a single Newton step is insufficient even for removing a single data point, and that at least two steps are required to ensure both privacy and accuracy. Our result leads us to conclude that the discrepancy in the number of steps arises because of the sub optimality of the notion of $\varepsilon$-certifiability and its incompatibility with noise adding mechanisms, which $\varepsilon$-Gaussian certifiability is able to overcome optimally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。