通过谱显著性筛选关键方向,实现高效且精准的机器遗忘。
Spectral Saliency for Machine Unlearning
- 基于谱显著性阈值筛选弱奇异分量,仅更新可信遗忘信号方向。
- 在图像分类、扩散模型和大语言模型上均实现有效遗忘,保留模型性能。
- 理论证明阈值策略平衡遗忘与保留,适合需要数据可控移除的场景。
机器遗忘(MU)旨在消除特定训练数据的影响,同时保持模型可用性。如其名称所示,MU可视为学习的逆过程,通过基于梯度的更新来抵消先前学习行为的影响。近期提出的Muon是一种梯度下降变体,采用谱幅度归一化以促进对稀有方向的探索,并表现出良好性能。受Muon启发,我们从谱视角出发,提出谱显著性遗忘(SSU)。SSU对弱奇异成分进行阈值处理,仅更新由强遗忘信号支持的方向。我们进一步从遗忘-保留权衡的角度为该阈值方法提供理论支持。在图像分类器、扩散模型和大语言模型上的实验表明,SSU具有有效性。
原文摘要 · Abstract (English)
Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can be viewed as the inverse of learning, using gradient-based updates to reduce the influence of a forget-set by counteracting the previously learned behavior. Recently, Muon, a gradient descent variant, has been introduced. Muon applies spectral magnitude normalization to encourage exploration of rare directions and demonstrates promising performance. Inspired by Muon, we adopt the spectral view for unlearning and propose Spectral Saliency Unlearning (SSU). SSU thresholds weak singular components and updates only those directions supported by a confident unlearning signal. We further provide theoretical justification for this thresholding approach from the perspective of the forgetting-retention trade-off. Experiments across image classifiers, diffusion models, and LLMs demonstrate SSU's effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。