用公开数据降低删除模型时的性能损失,实现高效隐私保护
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data

- 引入不对称朗之万删除框架,利用公开数据降低隐私成本
- 公开数据量越大,所需噪声越小,性能损失显著减少
- 适合大规模数据删除场景,尤其在数据分布不同时仍有效
基于噪声的可认证机器删除目前面临硬性瓶颈:为确保删除安全性所需的噪声强度通常会严重损害模型性能,尤其是在大规模删除请求下。尽管利用公开数据是差分隐私中缓解此矛盾的标准方法,但在删除领域尚未被探索。本文提出不对称朗之万删除(ALU)框架,通过注入公开数据来缓解隐私成本。理论证明,公开数据注入可将删除成本降低至 $O(1/n_{\mathrm{pub}}^2)$ 阶,带来严格于重训练的计算优势。这确立了新控制机制:通过增加公开数据量,可减少高噪声需求及其带来的性能损失。关键地,我们分析了真实场景下的分布偏移问题,明确刻画了公共与私有数据源之间的差异对性能的影响。实验表明,ALU可在恒定数据比例的大规模删除中保持高实用性,而传统对称方法在此场景下难以适用。使用变分瑞尼散度和成员推断攻击的实证评估证实,ALU在合理分布偏移下仍能有效抵御隐私攻击并保留模型性能。
原文摘要 · Abstract (English)
Noise-based certified machine unlearning currently faces a hard ceiling: the noise magnitude required to certify unlearning typically destroys model utility, particularly for large-scale deletion requests. While leveraging public data is a standard technique in differential privacy to relax this tension, its role in unlearning remains unexplored. We address this gap by introducing Asymmetric Langevin Unlearning (ALU), a framework that uses public data to mitigate privacy costs. We prove that public data injection suppresses the unlearning cost by a factor of $O(1/n_{\mathrm{pub}}^2)$, guaranteeing a strict computational advantage over retraining. This establishes a new control mechanism: practitioners can mitigate the need for high noise-and the associated utility loss-by increasing the volume of public data. Crucially, we analyze the realistic setting of distribution mismatch, explicitly characterizing how shifts between public and private sources impact utility. We show that ALU enables mass unlearning of constant dataset fractions -- a regime where standard symmetric methods become impractical -- while maintaining high utility. Empirical evaluations using variational Rényi divergence and membership inference attacks confirm that ALU effectively thwarts privacy attacks while preserving utility under reasonable distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。