arXiv:2601.04600cs.CLcs.AI2026-01

Rank-One编辑在多跳问答中效果差,新方法提升准确率96%

On the Limitations of Rank-One Model Editing in Answering Multi-hop Questions

  • 通过冗余编辑增强中间层知识传递,解决信息滞后问题
  • 2跳问答准确率提升15.5个百分点,较原方法提高96%
  • 适合需链式推理的模型知识更新场景

近期知识编辑(KE)进展,尤其是单阶知识编辑方法 Rank-One Model Editing(ROME),在更新Transformer中的单跳事实方面比微调和上下文学习更高效。然而,该方法在需要知识链式推理的多跳任务中面临显著挑战。本文研究了使用ROME编辑不同网络层深度的影响,识别出三种关键失败模式:第一,“跳得太晚”问题,即后层无法获取必要中间表示;第二,编辑后层时泛化能力急剧下降;第三,模型对编辑知识过度拟合,忽略上下文而优先选择已编辑的跳跃答案。为缓解“跳得太晚”和泛化能力衰减问题,我们提出冗余编辑(Redundant Editing)策略,简单有效。实验表明,该方法可使2跳问答准确率至少提升15.5个百分点,相较单次编辑策略提升96%,代价是部分特异性与语言自然性下降。

原文摘要 · Abstract (English)

Recent advances in Knowledge Editing (KE), particularly Rank-One Model Editing (ROME), show superior efficiency over fine-tuning and in-context learning for updating single-hop facts in transformers. However, these methods face significant challenges when applied to multi-hop reasoning tasks requiring knowledge chaining. In this work, we study the effect of editing knowledge with ROME on different layer depths and identify three key failure modes. First, the "hopping-too-late" problem occurs as later layers lack access to necessary intermediate representations. Second, generalization ability deteriorates sharply when editing later layers. Third, the model overfits to edited knowledge, incorrectly prioritizing edited-hop answers regardless of context. To mitigate the issues of "hopping-too-late" and generalisation decay, we propose Redundant Editing, a simple yet effective strategy that enhances multi-hop reasoning. Our experiments demonstrate that this approach can improve accuracy on 2-hop questions by at least 15.5 percentage points, representing a 96% increase over the previous single-edit strategy, while trading off some specificity and language naturalness.

知识编辑多跳推理模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。