arXiv:2602.01150cs.LGcs.AI2026-02被引 1

提出无需训练的统计会员推断方法,更可靠地检测模型是否真正遗忘数据。

SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing

  • 通过估计特征分布中非成员比例,无需训练影子模型即可审计遗忘效果
  • 实验证明其性能优于所有基于MIA的基线方法,且无需额外计算开销
  • 提供置信区间,可量化审计结果的可靠性,适合安全敏感场景

机器遗忘对于实现机器学习系统中的被遗忘权至关重要。其核心挑战在于如何可靠地审计模型是否真正遗忘指定训练数据。目前广泛采用会员推断攻击(MIA)进行遗忘审计,将无法被识别为成员的样本视为已成功遗忘。我们指出这一假设存在根本性缺陷:未能检测到会员身份并不等于真正遗忘。我们证明,被遗忘样本在特征空间中的位置与非成员样本存在本质差异,这种对齐偏差不可避免且不可观测,导致对遗忘性能的评估系统性偏乐观。同时,训练影子模型进行MIA带来巨大计算开销。为此,我们提出统计会员推断(SMI),一种无需训练的审计框架,将审计问题重新表述为估计未学习特征分布中非成员混合比例。除估算遗忘率外,SMI还提供自助法参考范围,用于量化审计可靠性。大量实验表明,SMI在无需影子模型训练的情况下,持续优于所有基于MIA的基线方法。总体而言,SMI为基于MIA的审计方法提供了具有理论保证和强实证表现的原理性、高效替代方案。

原文摘要 · Abstract (English)

Machine unlearning (MU) is essential for enforcing the right to be forgotten in machine learning systems. A key challenge of MU is how to reliably audit whether a model has truly forgotten specified training data. Membership Inference Attacks (MIAs) are widely used for unlearned model auditing, where samples that evade membership detection are regarded as successfully forgotten. We show this assumption is fundamentally flawed: failed membership inference does not imply true forgetting. We prove that unlearned samples occupy fundamentally different positions in the feature space than non-member samples, making this alignment bias unavoidable and unobservable, which leads to systematically optimistic evaluations of unlearning performance. Meanwhile, training shadow models for MIA incurs substantial computational overhead. To address both limitations, we propose Statistical Membership Inference (SMI), a training-free auditing framework that reformulates auditing as estimating the non-member mixture proportion in the unlearned feature distribution. Beyond estimating the forgetting rate, SMI also provides bootstrap reference ranges for quantified auditing reliability. Extensive experiments show that SMI consistently outperforms all MIA-based baselines, with no shadow model training required. Overall, SMI establishes a principled and efficient alternative to MIA-based auditing methods, with both theoretical guarantees and strong empirical performance.

机器遗忘会员推断审计统计方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。