arXiv:2505.18783cs.LGcs.AI2025-05

提出软加权机制,缓解模型删训中的过度遗忘问题。

Soft Weighted Machine Unlearning

  • 用凸二次规划求解样本权重,实现细粒度模型调整
  • 在公平性和鲁棒性任务中显著提升性能且减少效用下降
  • 可无缝集成到多数现有删训算法,适合需要精准修正的场景

机器删训作为一种事后处理技术,被广泛用于解决偏见缓解和鲁棒性增强等挑战,即为公平性和鲁棒性而进行的删训。然而,现有非隐私驱动的删训方法仍沿用专为隐私保护设计的二值数据移除框架,导致显著的信息损失,这种现象被称为过度删训。尽管过度删训常被认为主要造成性能下降,本文通过反事实留一分析深入探究其根本原因。我们提出一种加权影响函数,通过解析求解凸二次规划问题为每个样本分配定制化权重。基于此,提出软加权框架,实现对模型的精细调整以应对过度删训问题。实验表明,该软加权方案具有通用性,可无缝融入大多数现有删训算法,在公平性和鲁棒性任务中显著优于硬加权方案,在公平性/鲁棒性指标上表现更优,同时缓解效用指标下降,使删训算法成为有效的修正工具。

原文摘要 · Abstract (English)

Machine unlearning, as a post-hoc processing technique, has gained widespread adoption in addressing challenges like bias mitigation and robustness enhancement, colloquially, machine unlearning for fairness and robustness. However, existing non-privacy unlearning-based solutions persist in using binary data removal framework designed for privacy-driven motivation, leading to significant information loss, a phenomenon known as over-unlearning. While over-unlearning has been largely described in many studies as primarily causing utility degradation, we investigate its fundamental causes and provide deeper insights in this work through counterfactual leave-one-out analysis. In this paper, we introduce a weighted influence function that assigns tailored weights to each sample by solving a convex quadratic programming problem analytically. Building on this, we propose a soft-weighted framework enabling fine-grained model adjustments to address the over-unlearning challenge. We demonstrate that the proposed soft-weighted scheme is versatile and can be seamlessly integrated into most existing unlearning algorithms. Extensive experiments show that in fairness- and robustness-driven tasks, the soft-weighted scheme significantly outperforms hard-weighted schemes in fairness/robustness metrics and alleviates the decline in utility metric, thereby enhancing machine unlearning algorithm as an effective correction solution.

机器删训公平性鲁棒性加权机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。