arXiv:2507.05937cs.CL2025-07中稿 · ACL被引 1

揭示知识编辑评估方法的偏差,提醒别被表面指标误导。

Towards a Principled Evaluation of Knowledge Editors

  • 对比多种评估方法发现排名差异显著
  • 不同编辑方式在通用语言理解任务中表现不一
  • 指出常用字符串匹配法易产生假阳性结果

近年来,模型编辑受到越来越多关注,尤其是知识编辑领域。尽管近期发布了更具挑战性的评估数据集,但这些数据集采用不同方法评分编辑成功率,其稳健性与公平性仍待深入探究。我们发现,评估指标、方法及编辑批量大小的选择会显著改变知识编辑器的排名。关键的是,这种影响也体现在与知识编辑并行的通用语言理解任务上。此外,我们对当前主流数据集中偏好的基于字符串匹配的评估方法进行了人工分析,发现其存在产生假阳性匹配的倾向。

原文摘要 · Abstract (English)

Model editing has been gaining increasing attention over the past few years. For Knowledge Editing in particular, more challenging evaluation datasets have recently been released. These datasets use different methodologies to score the success of editors. Yet, it remains under-explored how robust these methodologies are and whether they unfairly favor some editors. Moreover, the disruptive impact of these editors on overall model capabilities remains a constant blind spot. We address both of these problems and show that choosing different metrics and evaluation methodologies as well as different edit batch sizes can lead to a different ranking of knowledge editors. Crucially we demonstrate this effect also on general language understanding tasks evaluated alongside the knowledge editing tasks. Further we include a manual assessment of the string matching based evaluation method for knowledge editing that is favored by recently released datasets, revealing a tendency to produce false positive matches.

知识编辑评估方法模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。