arXiv:2502.19770cs.CRcs.LG2025-02被引 12

无需训练参与,可高效审计模型删数据效果

TAPE: Tailored Posterior Difference for Auditing of Machine Unlearning

  • 通过影子模型快速模拟删数据后的模型差异
  • 重建私有信息验证删除效果,支持多样本请求
  • 比现有方法快4.5倍,适合真实场景审计

随着网络平台处理海量用户数据,机器遗忘成为保障用户被遗忘权的关键机制,允许用户请求从训练模型中移除其指定数据。然而,机器遗忘的审计仍严重缺乏研究。现有基于后门的方法效率低且不实用,需在初始训练阶段植入后门。本文提出一种独立于原始训练过程的审计方法TAPE(Tailored Posterior Difference)。我们发现机器遗忘过程会引入模型变化,其中包含被删除数据的信息。TAPE利用遗忘前后模型差异来评估信息清除程度。首先,通过一阶影响估计快速构建遗忘影子模型;其次,训练重构器模型提取并评估遗忘后验差异中的私有信息以实现审计。针对现有方法仅适用于单样本更新的问题,我们提出未学习数据扰动和基于影响的划分策略,提升多样本遗忘场景下的重构有效性。大量实验表明,TAPE在性能上显著优于现有最优方法,效率至少提升4.5倍,并支持更广泛的遗忘场景。

原文摘要 · Abstract (English)

With the increasing prevalence of Web-based platforms handling vast amounts of user data, machine unlearning has emerged as a crucial mechanism to uphold users' right to be forgotten, enabling individuals to request the removal of their specified data from trained models. However, the auditing of machine unlearning processes remains significantly underexplored. Although some existing methods offer unlearning auditing by leveraging backdoors, these backdoor-based approaches are inefficient and impractical, as they necessitate involvement in the initial model training process to embed the backdoors. In this paper, we propose a TAilored Posterior diffErence (TAPE) method to provide unlearning auditing independently of original model training. We observe that the process of machine unlearning inherently introduces changes in the model, which contains information related to the erased data. TAPE leverages unlearning model differences to assess how much information has been removed through the unlearning operation. Firstly, TAPE mimics the unlearned posterior differences by quickly building unlearned shadow models based on first-order influence estimation. Secondly, we train a Reconstructor model to extract and evaluate the private information of the unlearned posterior differences to audit unlearning. Existing privacy reconstructing methods based on posterior differences are only feasible for model updates of a single sample. To enable the reconstruction effective for multi-sample unlearning requests, we propose two strategies, unlearned data perturbation and unlearned influence-based division, to augment the posterior difference. Extensive experimental results indicate the significant superiority of TAPE over the state-of-the-art unlearning verification methods, at least 4.5$\times$ efficiency speedup and supporting the auditing for broader unlearning scenarios.

机器遗忘隐私审计后验差异数据删除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。