arXiv:2505.23291cs.CL2025-05ACL被引 3

构建剧本式评测框架,检验知识编辑在真实场景下的效果。

ScEdit: Script-based Assessment of Knowledge Editing

  • 设计剧本驱动的评测任务,涵盖反事实与时间类编辑
  • 所有方法在文本级指标上表现下降,验证任务挑战性
  • 支持动作型问题评估,适合大模型智能体研究者

知识编辑(KE)受到越来越多关注,但现有任务仍较简单。当前评估框架下,许多编辑方法得分极高,接近完美。然而,极少研究将KE融入真实应用场景(如大模型代理)。为此,我们提出新基准ScEdit(剧本式知识编辑基准),包含反事实和时间类编辑。结合词粒度与文本粒度评估方法,全面分析现有KE技术。该基准将传统基于事实的“是什么”问答扩展为基于行动的“如何做”问答。结果发现,所有方法在标准指标上性能下降,在文本级指标上面临挑战,表明任务具有高难度。基准代码已开源:https://github.com/asdfo123/ScEdit。

原文摘要 · Abstract (English)

Knowledge Editing (KE) has gained increasing attention, yet current KE tasks remain relatively simple. Under current evaluation frameworks, many editing methods achieve exceptionally high scores, sometimes nearing perfection. However, few studies integrate KE into real-world application scenarios (e.g., recent interest in LLM-as-agent). To support our analysis, we introduce a novel script-based benchmark -- ScEdit (Script-based Knowledge Editing Benchmark) -- which encompasses both counterfactual and temporal edits. We integrate token-level and text-level evaluation methods, comprehensively analyzing existing KE techniques. The benchmark extends traditional fact-based ("What"-type question) evaluation to action-based ("How"-type question) evaluation. We observe that all KE methods exhibit a drop in performance on established metrics and face challenges on text-level metrics, indicating a challenging task. Our benchmark is available at https://github.com/asdfo123/ScEdit.

知识编辑评估基准大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。