arXiv:2511.04473cs.LGcs.AI2025-11被引 2

构建真实图谱答案数据集,提升大模型知识检索的训练与评估效果

Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs

  • 用大模型自动生成带真实图谱答案的问答数据
  • 在未见过的图结构和关系类型上验证模型零样本泛化能力
  • 为知识图谱增强型大模型提供更可靠训练与评测基准

从图结构知识库中检索信息是提升大模型事实性的重要方向。现有方法缺乏具备真实答案目标的挑战性问答数据集,难以进行有效比较。本文提出 SynthKGQA,一种基于大模型的框架,可从任意知识图谱生成高质量的知识图谱问答数据集,并提供完整的图谱事实作为推理依据。该数据集不仅支持对知识图谱检索器的更精准评估,还能用于训练性能更优的模型。我们以 Wikidata 为基础生成 GTSQA,一个专为测试知识图谱检索器在未见图结构与关系类型上的零样本泛化能力而设计的新数据集,并在此上基准化多种主流知识图谱增强型大模型方案。

原文摘要 · Abstract (English)

Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various solutions have been proposed, a comparison of methods is difficult due to the lack of challenging QA datasets with ground-truth targets for graph retrieval. We present SynthKGQA, an LLM-powered framework for generating high-quality Knowledge Graph Question Answering datasets from any Knowledge Graph, providing the full set of ground-truth facts in the KG to reason over questions. We show how, in addition to enabling more informative benchmarking of KG retrievers, the data produced with SynthKGQA also allows us to train better models.We apply SynthKGQA to Wikidata to generate GTSQA, a new dataset designed to test zero-shot generalization abilities of KG retrievers with respect to unseen graph structures and relation types, and benchmark popular solutions for KG-augmented LLMs on it.

知识图谱大模型数据生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。