arXiv:2510.00625cs.AI2025-10被引 6

模型编辑看似成功,实则依赖隐藏捷径,真实效果存疑。

Is Model Editing Built on Sand? Revealing Its Illusory Success and Fragile Foundation

  • 通过少量参数修改实现知识更新,但易依赖表面捷径而非深层语义。
  • 在否定性查询下,顶尖编辑方法性能骤降,表明其泛化能力极弱。
  • 现有评估框架缺负例设计,导致结果虚假可信,需重构评测体系。

大语言模型不可避免地包含过时或错误的知识。更新、删除或遗忘这些知识对对齐、安全等问题至关重要。为此,模型编辑作为一种有前景的范式应运而生:仅修改少量参数,使特定事实得到更新,同时保留其他知识。尽管此前研究宣称取得显著成功,我们发现其可靠性建立在脆弱基础上,当前文献很大程度上由虚假的成功驱动。将模型输出引导至目标且改动最小的根本目标,会诱使模型利用隐藏捷径,而非真正理解语义。这一问题直接挑战了现有模型编辑研究的基础可行性,因为捷径与稳健的知识整合本质矛盾。巧合的是,这一问题长期被缺乏负例设计的评估框架所掩盖。为揭示它,我们系统构建了一套新评估方法。令人震惊的是,最先进的方法在最简单的否定查询下即告崩溃。实证证据表明,编辑行为很可能基于捷径而非完整语义,呼吁对模型编辑的根基进行紧急重新审视,否则后续发展难以有意义推进。

原文摘要 · Abstract (English)

Large language models (LLMs) inevitably encode outdated or incorrect knowledge. Updating, deleting, and forgetting such knowledge is important for alignment, safety, and other issues. To address this issue, model editing has emerged as a promising paradigm: by precisely editing a small subset of parameters such that a specific fact is updated while preserving other knowledge. Despite its great success reported in previous papers, we find the apparent reliability of editing rests on a fragile foundation and the current literature is largely driven by illusory success. The fundamental goal of steering the model's output toward a target with minimal modification would encourage exploiting hidden shortcuts, rather than utilizing real semantics. This problem directly challenges the feasibility of the current model editing literature at its very foundation, as shortcuts are inherently at odds with robust knowledge integration. Coincidentally, this issue has long been obscured by evaluation frameworks that lack the design of negative examples. To uncover it, we systematically develop a suite of new evaluation methods. Strikingly, we find that state-of-the-art approaches collapse even under the simplest negation queries. Our empirical evidence shows that editing is likely to be based on shortcuts rather than full semantics, calling for an urgent reconsideration of the very basis of model editing before further advancements can be meaningfully pursued.

模型编辑大模型评估漏洞语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。