arXiv:2505.23026cs.CLcs.AI2025-05ACL被引 6

提升大模型知识编辑在真实语境下的稳定性

Context-Robust Knowledge Editing for Language Models

  • 通过最小化隐藏状态对上下文的敏感性增强编辑鲁棒性
  • 在有前置上下文时编辑成功率显著提升
  • 适用于需要可靠知识更新的对话系统场景

知识编辑(KE)方法为修改大语言模型中的知识提供了高效途径。当前的评估通常仅关注编辑内容本身,忽略前后文的影响。但在实际应用中,前置上下文常触发原始知识,导致编辑失效。为此,我们构建了CHED基准,用于评估KE方法的上下文鲁棒性。评测结果显示,多数方法在存在前置上下文时表现不佳。为此,我们提出CoRE方法,通过最小化模型对编辑知识的隐藏状态上下文敏感性,增强其鲁棒性。该方法不仅在含前置上下文场景中提升了编辑成功率,还保持了模型整体能力。我们深入分析了用户话语与助手回复作为前置上下文的不同影响,并解析注意力得分模式,揭示特定词元如何影响编辑效果。

原文摘要 · Abstract (English)

Knowledge editing (KE) methods offer an efficient way to modify knowledge in large language models. Current KE evaluations typically assess editing success by considering only the edited knowledge without any preceding contexts. In real-world applications, however, preceding contexts often trigger the retrieval of the original knowledge and undermine the intended edit. To address this issue, we develop CHED -- a benchmark designed to evaluate the context robustness of KE methods. Evaluations on CHED show that they often fail when preceding contexts are present. To mitigate this shortcoming, we introduce CoRE, a KE method designed to strengthen context robustness by minimizing context-sensitive variance in hidden states of the model for edited knowledge. This method not only improves the editing success rate in situations where a preceding context is present but also preserves the overall capabilities of the model. We provide an in-depth analysis of the differing impacts of preceding contexts when introduced as user utterances versus assistant responses, and we dissect attention-score patterns to assess how specific tokens influence editing success.

知识编辑大模型上下文鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。