高维场景下用两步牛顿法加噪声,实现数据删除的可认证安全
Certified Data Removal Under High-dimensional Settings
- 基于牛顿法迭代更新参数,结合缩放拉普拉斯噪声消除残留影响
- 高维下单步牛顿无效,需两步才能保证数据删除的可认证性
- 适合关注模型隐私保护与可证明安全性的研究人员
机器遗忘旨在高效移除训练数据对已训练模型的影响,无需重新训练。尽管低维情形(参数量p远小于样本量n)已有进展,但高维情形下理论保障仍具挑战。本文提出一种从原始参数出发、执行至多两步理论引导的牛顿步的遗忘算法。随后添加经精心缩放的各向同性拉普拉斯噪声,以彻底消除遗忘数据的潜在影响。当n, p趋于无穷且n/p为固定比值时,模型复杂度与有限信噪比之间的相互作用导致显著理论与计算障碍。数值实验验证了该方法在可认证性与精度上的有效性。
原文摘要 · Abstract (English)
Machine unlearning focuses on the computationally efficient removal of specific training data from trained models, ensuring that the influence of forgotten data is effectively eliminated without the need for full retraining. Despite advances in low-dimensional settings, where the number of parameters \( p \) is much smaller than the sample size \( n \), extending similar theoretical guarantees to high-dimensional regimes remains challenging. We propose an unlearning algorithm that starts from the original model parameters and performs a theory-guided sequence of Newton steps \( T \in \{ 1,2\}\). After this update, carefully scaled isotropic Laplacian noise is added to the estimate to ensure that any (potential) residual influence of forget data is completely removed. We show that when both \( n, p \to \infty \) with a fixed ratio \( n/p \), significant theoretical and computational obstacles arise due to the interplay between the complexity of the model and the finite signal-to-noise ratio. Finally, we show that, unlike in low-dimensional settings, a single Newton step is insufficient for effective unlearning in high-dimensional problems -- however, two steps are enough to achieve the desired certifiebility. We provide numerical experiments to support the certifiability and accuracy claims of this approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。