arXiv:2601.09624cs.LGcs.AI2026-01ACL被引 3

提出可预测遗忘难易度的电路级指标,揭示模型内部记忆保护机制。

A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning

  • 基于模型电路结构设计预遗忘难度评分,量化每条样本的遗忘难易程度。
  • 简单样本路径短且集中于早期层,困难样本依赖深层复杂路径。
  • 指标稳定且可解释,助力开发基于机制的高效遗忘方法。

机器遗忘对构建可信合规的语言模型至关重要。然而,相同方法下不同样本的遗忘效果差异显著:部分样本能可靠擦除,而另一些则顽固留存。我们指出,这种差异不仅是数据层面的现象,更反映了模型内部编码与保护记忆的机制。本文从机制视角出发,基于模型电路——决定预测生成的结构化交互路径——提出电路引导的遗忘难度(CUD)指标,该指标在遗忘前为每个样本分配连续难度分数。大量实验表明,CUD能可靠区分内在易忘与难忘样本,且在不同遗忘方法间保持稳定。我们识别出关键电路模式:易忘样本关联较短、较浅的交互路径,集中在原模型的早期至中期;而难忘样本依赖更长更深的路径,靠近晚期计算阶段。相比现有定性研究,CUD首次实现对遗忘难度的原理性、细粒度和可解释分析,推动基于模型机制的遗忘方法发展。

原文摘要 · Abstract (English)

Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We argue that this disparity is not only a data-side phenomenon, but also reflects model-internal mechanisms that encode and protect memorized information. We study this problem from a mechanistic perspective based on model circuits--structured interaction pathways that govern how predictions are formed. We propose Circuit-guided Unlearning Difficulty (CUD), a {\em pre-unlearning} metric that assigns each sample a continuous difficulty score using circuit-level signals. Extensive experiments demonstrate that CUD reliably separates intrinsically easy and hard samples, and remains stable across unlearning methods. We identify key circuit-level patterns that reveal a mechanistic signature of difficulty: easy-to-unlearn samples are associated with shorter, shallower interactions concentrated in earlier-to-intermediate parts of the original model, whereas hard samples rely on longer and deeper pathways closer to late-stage computation. Compared to existing qualitative studies, CUD takes a first step toward a principled, fine-grained, and interpretable analysis of unlearning difficulty; and motivates the development of unlearning methods grounded in model mechanisms.

机器遗忘模型机制可解释性电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。