arXiv:2609.02285cs.LGcs.CL2026-09

实验证明模型知识越解耦,删除信息时副作用越小。

Entangled Representations Amplify Collateral Damage in Unlearning

  • 通过控制模型解耦程度,测试知识纠缠对遗忘的影响。
  • 解耦模型在遗忘相同内容时,保留代价低至1/4(两方法)或1/1.3(一方法)。
  • 为可解释性研究中长期猜想提供直接证据,适合关注隐私与模型安全的研究者。

可解释性研究普遍认为,神经网络中不同知识域之间的表征纠缠会增加遗忘难度。然而这一直未在受控实验中得到验证。本文通过复用选择性梯度掩码(SGTM),训练了六组参数量为254M的英文语言模型,其生物学与非生物学知识的解耦程度逐级变化。对每组模型应用三种标准遗忘方法,结果表明:解耦程度更高的模型始终具有更优的保留-遗忘权衡——在固定遗忘水平下,其中两种方法的保留代价降低约4倍,第三种方法降低1.3倍。由于仅改变模型结构而保持数据和遗忘算法不变,该结果直接证明表征纠缠是遗忘过程中副作用(即‘附带损伤’)的重要成因,支持了长期存在的理论猜想。类似实验设计也可用于检验其他可解释性假设。

原文摘要 · Abstract (English)

A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly tested in a controlled experiment. We present a way to do so: by repurposing Selective Gradient Masking (SGTM), we train a suite of six 254M-parameter language models on English Wikipedia with graded levels of disentanglement between biology and non-biology knowledge. Applying three standard unlearning methods to every model in the suite, we find that more disentangled models consistently achieve better retain-forget trade-offs: at a fixed level of forgetting, the most disentangled models incur roughly $4\times$ lower retain cost under two of the three methods, and $1.3\times$ lower under the third. Because our intervention changes only the model, not the data or the unlearning algorithm, this is direct evidence that representational entanglement is one of the causes of collateral damage in unlearning, as interpretability researchers have long suspected. A similar design could be used to test other structural claims from interpretability.

模型遗忘表征解耦可解释性隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。