arXiv:2506.00829cs.CLcs.AI2025-06ACL被引 7

新基准COMPKE评估知识编辑后模型应对复杂问题的能力

COMPKE: Complex Question Answering under Knowledge Editing

  • 构建包含11924个真实场景复杂问题的新评测集
  • 发现不同编辑方法在不同模型上效果差异巨大,最高差超10倍
  • 适合关注大模型知识更新可靠性与泛化能力的研究者

知识编辑能高效修改大语言模型中的知识,受到广泛关注。现有基准主要使用多跳问答评估新注入或更新的知识,但我们认为这些基准未能有效检验模型在真实场景中应用更新知识的能力,尤其在涉及一对多关系或多步逻辑推理的复杂问题上。为此,我们提出新基准COMPKE:知识编辑下的复杂问题问答,包含11,924个反映真实情境的复杂问题。我们在COMPKE上对四种知识编辑方法进行广泛评估,发现其效果在不同模型间差异显著。例如,MeLLo在GPT-4O-MINI上准确率达39.47,但在QWEN2.5-3B上骤降至3.83。我们从方法和模型两个角度深入分析了这种差异的成因。数据集已开源于https://github.com/kzjkzj666/CompKE。

原文摘要 · Abstract (English)

Knowledge Editing, which efficiently modifies the knowledge in large language models, has gathered great attention. Current benchmarks primarily use multi-hop question answering to assess and analyze newly injected or updated knowledge. However, we argue that these benchmarks fail to effectively evaluate how well the updated models apply this knowledge in real-life scenarios, particularly when questions require complex reasoning, involving one-to-many relationships or multi-step logical intersections. To fill in this gap, we introduce a new benchmark, COMPKE: Complex Question Answering under Knowledge Editing, which includes 11,924 complex questions that reflect real-life situations. We conduct an extensive evaluation of four knowledge editing methods on COMPKE, revealing that their effectiveness varies notably across different models. For instance, MeLLo attains an accuracy of 39.47 on GPT-4O-MINI, but this drops sharply to 3.83 on QWEN2.5-3B. We further investigate the underlying causes of these disparities from both methodological and model-specific perspectives. The datasets are available at https://github.com/kzjkzj666/CompKE.

知识编辑复杂推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。