arXiv:2606.18309cs.LGcs.AI2026-06

提出SAGE方法,用后处理修复大模型遗忘时的保留能力损失。

SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector

论文配图:SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector
图 1 · 摘自论文原文
  • 基于激活几何结构的后处理修正最终更新向量
  • 在多方法、多规模下显著缓解遗忘与保留的权衡
  • 无需重跑原流程,适合快速修复遗忘模型

大型语言模型遗忘旨在移除不良知识或行为的同时保留原有能力。现有方法均存在遗忘与保留之间的权衡。我们发现保留激活偏差可量化任何遗忘方法对保留性能的损伤,而无需考虑具体实现。由此可采用后处理方式恢复保留能力。为此,我们提出互补的后处理框架,通过SAGE(Spectral Activation-GEometry Sanitization)对最终更新向量进行无源修正。SAGE从少量保留代理中获取真实模块输入,提取其主导激活几何结构,在闭式解中求解源锚定优化目标,抑制与高能量保留方向对齐的更新分量,同时保留原始方法的遗忘载体。在多种遗忘方法、模型规模和基准测试中,SAGE始终有效缓解保留-遗忘权衡,揭示了对最终更新向量进行后处理是机器遗忘中一个实用且未被充分探索的方向。

原文摘要 · Abstract (English)

Large Language Model (LLM) unlearning aims to remove undesirable knowledge or behaviors while preserving retained capabilities. Current unlearning methods all involve a trade-off between unlearning and retention. We have found that the retention activation bias can also be used to quantify the damage an unlearning method inflicts on retention, without considering the specific implementation of the unlearning process. This allows us to restore retention performance for any unlearning method using a post-hoc approach. Therefore, we propose a complementary post-hoc setting to sanitize the final update vector without rerunning the original unlearning pipeline. In this setting, we design SAGE, Spectral Activation-GEometry Sanitization, a source-agnostic correction for final unlearning updates. SAGE collects real module inputs from a small retain proxy, extracts their dominant activation geometry, and solves a source-anchored optimization objective in closed form, which suppresses update components aligned with high-energy retained directions while preserving the source method's forgetting carrier. Across multiple unlearning methods, model scales, and benchmarks, SAGE consistently relieves the retain-forget trade-off, identifying post-hoc sanitization of final vectors as a practical and underexplored axis for machine unlearning.

模型遗忘后处理大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。