arXiv:2605.28839cs.LG2026-05ACL被引 2

发现知识编辑共用同一套权重机制,可被一个掩码反向删除。

One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

论文配图:One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
图 1 · 摘自论文原文
  • 通过训练二值掩码定位关键权重子集,实现对多种编辑的统一控制。
  • 掩码在训练集上逆转80%、测试集上逆转超70%的编辑结果。
  • 揭示编辑本质是抑制而非覆盖旧知识,解释失败传播原因。

知识编辑方法如ROME和MEMIT通过修改MLP权重来更新Transformer模型中的事实关联。尽管主要通过输出行为评估,其内部机制仍不明确。我们探究不同事实的编辑是否依赖共同机制。尽管权重变化因事实而异,我们提出ROME和MEMIT均针对同一组对维持编辑至关重要的权重。为隔离该子集,我们在编辑权重上训练一个紧凑的二值掩码。该掩码在训练集上逆转80%的编辑,在测试集上逆转超70%,证实多样编辑共享相同功能结构。分析显示,掩码通过消除后期层的过度关注来逆转编辑。此外,编辑时注入掩码使成功率从98%降至38%,表明该机制对编辑成功至关重要。发现编辑通过抑制而非覆盖知识,解释了为何ROME和MEMIT无法传播到相关事实。所识别的共同功能子空间为检测与防御非预期编辑提供依据。

原文摘要 · Abstract (English)

Knowledge editing methods such as ROME and MEMIT update factual associations in transformer models by modifying MLP weights. While evaluated mainly by output behavior, their internal mechanism remains underexplored. We investigate whether edits rely on a common mechanism, regardless of which fact is modified. Despite fact-specific weight changes, we argue that ROME and MEMIT target the same subset of weights critical for maintaining edits. To isolate this subset, we train a compact binary mask over the edited weights. The mask reverses 80% of edits on the training set and over 70% on the test set, confirming that diverse edits share a common functional structure. Our analysis reveals that the mask reverses edits by eliminating overattention in later layers. Additionally, we show that injecting the mask during editing drops editing success from 98% to 38%, demonstrating that this mechanism is necessary for edits to succeed. Our finding that edits suppress rather than overwrite knowledge explains why ROME and MEMIT fail to propagate changes to related facts. The identified common functional subspace informs detection and defense against unwanted edits.

知识编辑权重分析模型可解释性掩码机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。