用可解释框架揭示模型删知识后仍会残留的真相
Understanding Empirical Unlearning with Combinatorial Interpretability
- 用组合可解释性直接分析模型权重中的知识残留
- 发现多数方法只抑制表达,未真正删除底层信息
- 揭示知识在微调后易被恢复,适合安全与可信AI研究
尽管近期众多方法致力于从预训练模型中移除特定知识,但看似已被清除的知识仍可能以多种方式被恢复。由于大型基础模型缺乏可解释性,理解这些知识是否真实消失仍是重大挑战。为此,本文采用最近发展的组合可解释性框架,该框架专为两层神经网络设计,能够直接观察模型权重中编码的知识。我们在该框架下重现了基线去学习方法,并从两个维度考察其行为:(i)是否真正移除了目标概念的知识(我们希望删除的内容),还是仅抑制其表达而保留底层信息;(ii)被声称删除的知识在各种微调操作下有多容易被恢复。结果在完全可解释的设定下揭示了知识为何会在去学习后依然存续以及何时可能重现。
原文摘要 · Abstract (English)
While many recent methods aim to unlearn or remove knowledge from pretrained models, seemingly erased knowledge often persists and can be recovered in various ways. Because large foundation models are far from interpretable, understanding whether and how such knowledge persists remains a significant challenge. To address this, we turn to the recently developed framework of combinatorial interpretability. This framework, designed for two-layer neural networks, enables direct inspection of the knowledge encoded in the model weights. We reproduce baseline unlearning methods within the combinatorial interpretability setting and examine their behavior along two dimensions: (i) whether they truly remove knowledge of a target concept (the concept we wish to remove) or merely inhibit its expression while retaining the underlying information, and (ii) how easily the supposedly erased knowledge can be recovered through various fine-tuning operations. Our results shed light within a fully interpretable setting on how knowledge can persist despite unlearning and when it might resurface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。