arXiv:2606.10554cs.CLcs.AI2026-06中稿 · the 24th Internati…

测试大模型改知识后能否推导出逻辑结论,发现现有方法严重失效。

Benchmarking Knowledge Editing using Logical Rules

论文配图:Benchmarking Knowledge Editing using Logical Rules
图 1 · 摘自论文原文
  • 从知识图谱提取逻辑规则,生成多跳问题评估推理能力。
  • 主流方法在直接事实编辑上准确率高,但推导结论错误率超24%。
  • 适合关注知识推理与模型可信度的研究者和工程师。

大型语言模型在需实时更新知识的应用中日益重要,但重训练成本过高,因此知识编辑技术至关重要。现有评测主要关注编辑事实的召回率,常忽略其逻辑后果。为此,本文提出新基准,通过从知识图谱提取与编辑相关的逻辑规则,生成多跳问题以评估逻辑推导能力。实验表明,尽管如ROME和FT等方法能准确插入直接陈述,却难以注入蕴含知识,对直接知识的评估与蕴含知识的评估之间存在高达24%的性能差距。这凸显了语义感知评估框架在知识编辑中的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLMs is computationally expensive. Therefore, knowledge editing techniques are crucial for maintaining current information and correcting erroneous assertions within pre-trained models. Current benchmarks for knowledge editing primarily focus on recalling edited facts, often neglecting their logical consequences. To address this limitation, we introduce a new benchmark designed to evaluate how knowledge editing methods handle the logical consequences of a single fact edit. Our benchmark extracts relevant logical rules from a knowledge graph for a given edit. Then, it generates multi-hop questions based on these rules to assess the impact on logical consequences. Our findings indicate that while existing knowledge editing approaches can accurately insert direct assertions into LLMs, they frequently fail to inject entailed knowledge. Specifically, experiments with popular methods like ROME and FT reveal a substantial performance gap, up to 24%, between evaluations on directly edited knowledge and on entailed knowledge. This highlights the critical need for semantics-aware evaluation frameworks in knowledge editing.

知识编辑逻辑推理评测基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。