现有模型编辑方法在真实场景中效果远差于报告值,因评估方式存在严重偏差。
The Mirage of Model Editing: Revisiting Evaluation in the Wild
- 设计新基准QAEdit与无任务依赖评估框架WILD,更贴近真实应用
- 实测编辑准确率仅38.5%,远低于此前96.8%的报告结果
- 揭示测试时教师强制导致真值泄露,是性能虚高的主因
尽管文献中报告的编辑效果接近完美,但模型编辑在真实场景中的有效性仍不明确。为弥合这一差距,我们提出与主流问答数据集对齐的QAEdit基准,以及面向任务的无偏评估框架WILD。单次编辑实验显示,当前方法实际表现远低于以往报告(38.5% vs. 96.8%)。问题根源在于先前合成评估中存在的缺陷,其中最严重的是测试阶段使用教师强制,导致真值内容和长度泄露,造成性能高估。此外,通过模拟连续编辑部署,发现现有方法在仅1000次编辑后即严重失效。本工作呼吁研究界转向严谨评估,并发展可扩展、鲁棒的模型编辑方法,以支持大语言模型在真实场景中的知识更新。
原文摘要 · Abstract (English)
Despite near-perfect results reported in the literature, the effectiveness of model editing in real-world applications remains unclear. To bridge this gap, we introduce QAEdit, a new benchmark aligned with widely used question answering (QA) datasets, and WILD, a task-agnostic evaluation framework designed to better reflect real-world usage of model editing. Our single editing experiments show that current editing methods perform substantially worse than previously reported (38.5% vs. 96.8%). We demonstrate that it stems from issues in the synthetic evaluation practices of prior work. Among them, the most severe is the use of teacher forcing during testing, which leaks both content and length of the ground truth, leading to overestimated performance. Furthermore, we simulate practical deployment by sequential editing, revealing that current approaches fail drastically with only 1000 edits. This work calls for a shift in model editing research toward rigorous evaluation and the development of robust, scalable methods that can reliably update knowledge in LLMs for real-world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。