提出新方法让模型删除数据时更少影响其他数据性能。
Retain-Neutral Surrogates for Min-Max Unlearning

- 用正交约束构造保留中立的代理点,避免删除时误伤性能。
- 在高耦合场景下显著降低保留损失,实验验证优于传统方法。
- 适合需要精准删除数据且保持模型整体表现的应用场景。
机器遗忘旨在移除特定训练数据的影响,同时保持对剩余数据的性能。近似遗忘可视为局部编辑问题;在极小-极大遗忘中,关键局部对象是评估保留目标的代理点。当遗忘与保留梯度强烈对齐时,无约束的遗忘最大化扰动可能移动到使保留损失增加的代理点。本文提出保留正交代理遗忘(ROSU),通过在固定扰动预算下,最大化一阶遗忘增益并约束一阶保留变化为零,来限制内部代理构建。该方法产生闭式保留正交扰动、轻量级外层更新,并沿保留中立方向放大效果。理论分析表明:(i) 保留损伤具有曲率控制的二阶上界;(ii) 在正向对齐区域,ROSU严格优于标准极小-极大扰动;(iii) 当两梯度近乎正交时二者近似等价。在视觉与语言基准(CIFAR-10/100、Tiny-ImageNet、TOFU、WMDP)上的实证模式符合此几何特性:ROSU在高耦合场景下优势明显,其余场景也保持竞争力。
原文摘要 · Abstract (English)
Machine unlearning seeks to remove the influence of designated training data while preserving performance on the remaining data. Approximate unlearning can be viewed as a local editing problem; in min-max unlearning, the key local object is the surrogate point at which the retain objective is evaluated. When forget and retain gradients are strongly aligned, an unconstrained forget-maximizing perturbation can move to a surrogate point that increases retain loss. We propose Retain-Orthogonal Surrogate Unlearning (ROSU), which constrains the inner surrogate construction by maximizing first-order forget gain subject to zero first-order retain change under a fixed perturbation budget. This yields a closed-form retain-orthogonal perturbation, a lightweight transported outer update, and amplification along the retain-neutral direction. Our analysis establishes (i) a curvature-controlled second-order bound on retain damage, (ii) a positive-alignment regime in which ROSU strictly reduces surrogate retain loss relative to standard min-max perturbations, and (iii) near-equivalence when the two gradients are nearly orthogonal. Across vision and language benchmarks (CIFAR-10/100, Tiny-ImageNet, TOFU, WMDP), the empirical pattern follows this geometry: ROSU gives its clearest gains in high-coupling regimes while remaining competitive elsewhere.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。