用模式连通性解析机器遗忘机制,揭示遗忘路径与隐私保护的内在关联。
Understanding Machine Unlearning Through the Lens of Mode Connectivity
- 通过模式连通性分析遗忘过程中的参数路径平滑性。
- 发现遗忘模型常处于低损失连通基域,且遗忘非线性进行。
- 适合研究模型隐私、可解释性及抗重训练攻击的从业者。
机器遗忘旨在移除训练模型中的不期望信息,而无需从头重新训练。尽管近期取得进展,但遗忘过程中的损失景观与优化几何仍不清楚。本文通过模式连通性——即独立训练的模型可在参数空间中以低损失平滑路径相连——来研究机器遗忘。我们提出「遗忘中的模式连通性」(MCU),并在多种设置下评估,包括课程学习、二阶优化及不同遗忘方法间的连通性。结果表明,许多遗忘后的模型位于具有平滑保留/遗忘行为的连通基域中;训练动态变化可使解进入不同基域。MCU还揭示同一基域内模型在隐私指标上差异显著,且遗忘过程从原模型到遗忘模型呈非线性。此外,线性连通性表明多数近似遗忘方法在机理上不同于重新训练。最后,基于MCU的集成可提升泛化能力与抗重训练攻击鲁棒性,且MCU平滑度与遗忘难度相关。据我们所知,这是首个从模式连通性视角研究机器遗忘的工作。
原文摘要 · Abstract (English)
Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we study machine unlearning through the lens of mode connectivity--the phenomenon that independently trained models can often be connected by smooth low-loss paths in parameter space. We introduce {\em mode connectivity in unlearning} (MCU) and evaluate it across a range of settings, including curriculum learning, second-order optimization, and connectivity across different unlearning methods. We find that many unlearned models lie in connected basins with smooth retain/forget behavior, while changes in training dynamics can move solutions into different basins. MCU also reveals that models within the same basin can differ substantially on privacy metrics, and that unlearning progresses nonlinearly from the original model to the unlearned model. In addition, linear connectivity suggests that most approximate unlearning methods are mechanistically distinct from retraining. Finally, MCU-based ensembling can improve generalization and robustness to relearning attacks, and MCU smoothness correlates with unlearning difficulty. To our knowledge, this is the first study of machine unlearning through the lens of mode connectivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。