arXiv:2509.16525cs.SEcs.AI2025-09

用因果分析检测模型未彻底删数据,提升隐私与公平性

Causal Fuzzing for Verifying Machine Unlearning

  • 基于因果关系分析数据点和特征的删除影响
  • 在5个数据集上发现基线方法遗漏的残留影响
  • 适合关注模型隐私与可解释性的研究人员

随着机器学习模型广泛应用于决策系统,对特定数据或特征的'遗忘'能力对于提升模型适应性、公平性和隐私保护至关重要,尤其在训练成本高昂的情况下。为有效指导模型遗忘,充分验证不可或缺。现有验证方法洞察有限,常无法捕捉间接影响。本文提出CAFÉ,一种基于因果的统一框架,用于黑盒机器学习模型的数据点级与特征级遗忘验证。CAFÉ通过因果依赖关系评估遗忘目标的直接与间接影响,提供细粒度分析与可操作洞察。在五个数据集和三种模型架构上的评估表明,CAFÉ成功检测到基线方法遗漏的残留影响,同时保持计算效率。

原文摘要 · Abstract (English)

As machine learning models become increasingly embedded in decision-making systems, the ability to "unlearn" targeted data or features is crucial for enhancing model adaptability, fairness, and privacy in models which involves expensive training. To effectively guide machine unlearning, a thorough testing is essential. Existing methods for verification of machine unlearning provide limited insights, often failing in scenarios where the influence is indirect. In this work, we propose CAFÉ, a new causality based framework that unifies datapoint- and feature-level unlearning for verification of black-box ML models. CAFÉ evaluates both direct and indirect effects of unlearning targets through causal dependencies, providing actionable insights with fine-grained analysis. Our evaluation across five datasets and three model architectures demonstrates that CAFÉ successfully detects residual influence missed by baselines while maintaining computational efficiency.

模型遗忘因果推理隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。