检验局部参数更新能否真正实现知识遗忘,发现其效果并不如预期可靠。
Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
- 通过控制实验检验局部参数修改对遗忘的影响
- 发现需修改的参数集不固定,效果不稳定
- 挑战了局部更新能精准删除知识的假设
大语言模型常保留意外内容,促使知识遗忘研究兴起。近期方法强调局部化遗忘,仅更新特定参数区域以移除目标知识,同时保护其他通用知识。然而,由于缺乏对遗忘与知识保留之间权衡的严谨评估,其有效性尚不明确。本文重新审视现有局部遗忘方法,并通过受控实验严格检验局部参数更新是否因果性地促进遗忘。结果表明,有效遗忘所需的参数集合并非严格确定,挑战了局部化遗忘的核心假设——参数局部性可直接指示有效知识移除。
原文摘要 · Abstract (English)
Large language models often retain unintended content, prompting growing interest in knowledge unlearning. Recent approaches emphasize localized unlearning, restricting parameter updates to specific regions in an effort to remove target knowledge while preserving unrelated general knowledge. However, their effectiveness remains uncertain due to the lack of robust and thorough evaluation of the trade-off between the competing goals of unlearning. In this paper, we begin by revisiting existing localized unlearning approaches. We then conduct controlled experiments to rigorously evaluate whether local parameter updates causally contribute to unlearning. Our findings reveal that the set of parameters that must be modified for effective unlearning is not strictly determined, challenging the core assumption of localized unlearning that parameter locality is inherently indicative of effective knowledge removal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。