arXiv:2601.22028cs.LG2026-01被引 2

提出对比正则化方法,让大模型遗忘更彻底且不干扰保留知识。

From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning

  • 通过对比学习分离遗忘与保留特征,减少表征纠缠。
  • 在多个模型和数据集上显著降低遗忘-保留特征的纠缠度。
  • 无需额外隐私风险,适合需要安全遗忘的应用场景。

现有大模型遗忘方法多依赖预测空间中的对齐目标,虽能减少遗忘内容生成,但可能仅抑制输出,而使遗忘概念仍残留在表征中并纠缠于保留知识。本文提出CLReg,一种对比表征正则化方法,可识别遗忘特征并将其从保留特征中推开,从而显式降低遗忘-保留干扰,同时对保留特征扰动极小。我们首次提供理论分析,揭示表征塑造与纠缠减少的关系。在多个基准测试和不同规模的LLM上,CLReg有效降低了遗忘-保留表征纠缠,提升主流遗忘方法效果,且不引入额外隐私风险,为未来基于表征空间重塑的遗忘技术提供新思路。

原文摘要 · Abstract (English)

Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, explicitly reducing forget-retain interference with minimal shifts on retain features. We provide first theoretical insights that relate representation shaping to entanglement reduction. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement that facilitates mainstream unlearning methods without positing extra privacy risks, inspiring future work that reshapes the representation space to remove forget concepts.

大模型遗忘对比学习表征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。