arXiv:2508.08344cs.AI2025-08Conference of the …被引 8

测试知识图谱增强模型在信息缺失时的推理能力,发现多数依赖记忆而非真正推理。

What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge

  • 构建新评测基准,模拟知识缺失场景评估推理能力
  • 现有方法在知识不全时表现差,常靠模型记忆而非推理
  • 揭示不同设计对泛化能力影响,指导未来改进方向

基于知识图谱的检索增强生成(KG-RAG)正被广泛研究,旨在结合大语言模型的推理能力与知识图谱的结构化证据。然而当前评估方法存在缺陷:现有基准常包含可直接通过图谱三元组回答的问题,无法判断模型是推理还是简单检索。此外,评价指标不统一、答案匹配标准宽松,阻碍了有效比较。本文提出通用基准构建方法,发布BRINK(不完整知识下的推理评测基准),系统评估KG-RAG在知识缺失条件下的表现。实证结果表明,当前方法在知识不全时推理能力有限,多依赖内部记忆,且泛化性能因设计而异。

原文摘要 · Abstract (English)

Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) is an increasingly explored approach for combining the reasoning capabilities of large language models with the structured evidence of knowledge graphs. However, current evaluation practices fall short: existing benchmarks often include questions that can be directly answered using existing triples in KG, making it unclear whether models perform reasoning or simply retrieve answers directly. Moreover, inconsistent evaluation metrics and lenient answer matching criteria further obscure meaningful comparisons. In this work, we introduce a general method for constructing benchmarks and present BRINK (Benchmark for Reasoning under Incomplete Knowledge) to systematically assess KG-RAG methods under knowledge incompleteness. Our empirical results show that current KG-RAG methods have limited reasoning ability under missing knowledge, often rely on internal memorization, and exhibit varying degrees of generalization depending on their design.

知识图谱RAG推理评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。