arXiv:2505.23495cs.CLcs.AI2025-05NeurIPS被引 13

用AI修复知识图谱问答数据集缺陷,打造更可靠的评测基准

Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

  • 引入闭环框架自动识别并修正数据集中的错误和歧义问题
  • 构建1万条高质量问答数据,真实准确率仅57%的旧数据被大幅改善
  • 适合研究知识图谱推理与评估的学者,尤其关注模型真实能力验证

知识图谱问答(KGQA)系统依赖高质量基准测试来评估复杂的多跳推理能力。然而,尽管广为使用,WebQSP和CWQ等主流数据集存在严重质量问题,包括不准确或不完整的标注、模糊、简单或无法回答的问题,以及过时或不一致的知识。我们对16个常用KGQA数据集进行人工审核,发现平均事实正确率仅为57%。为此,我们提出KGQAGen——一个基于大模型的闭环生成框架,结合结构化知识锚定、大模型引导生成与符号化验证,生成可挑战且可验证的问答实例。利用该框架,我们构建了基于Wikidata的10,000条规模的基准数据集KGQAGen-10k,并评估多种KG-RAG模型。实验表明,即使最先进的系统在此基准上也表现不佳,凸显其揭示现有模型局限性的能力。研究呼吁更严格的基准构建标准,并将KGQAGen定位为推进KGQA评估的可扩展框架。

原文摘要 · Abstract (English)

Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critical quality issues, including inaccurate or incomplete ground-truth annotations, poorly constructed questions that are ambiguous, trivial, or unanswerable, and outdated or inconsistent knowledge. Through a manual audit of 16 popular KGQA datasets, including WebQSP and CWQ, we find that the average factual correctness rate is only 57 %. To address these issues, we introduce KGQAGen, an LLM-in-the-loop framework that systematically resolves these pitfalls. KGQAGen combines structured knowledge grounding, LLM-guided generation, and symbolic verification to produce challenging and verifiable QA instances. Using KGQAGen, we construct KGQAGen-10k, a ten-thousand scale benchmark grounded in Wikidata, and evaluate a diverse set of KG-RAG models. Experimental results demonstrate that even state-of-the-art systems struggle on this benchmark, highlighting its ability to expose limitations of existing models. Our findings advocate for more rigorous benchmark construction and position KGQAGen as a scalable framework for advancing KGQA evaluation.

知识图谱问答系统数据质量评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。