arXiv:2512.21799cs.IR2025-12被引 1

构建学术知识图谱与问答数据集,推动科研领域智能问答研究。

KG20C & KG20C-QA: Scholarly Knowledge Graph Benchmarks for Link Prediction and Question Answering

  • 从微软学术图谱筛选高质量论文构建知识图谱
  • 基于图谱生成自然语言问答对,支持多种模型评估
  • 提供可复现的评测协议,适合研究者做推理与问答实验

本文提出KG20C和KG20C-QA两个学术知识图谱基准数据集,用于推动学术领域问答(QA)研究。KG20C通过精选会议、质量过滤与模式定义,从微软学术图谱构建高质量学术知识图谱,此前虽在GitHub等非同行评审平台发布,但本论文首次提供正式同行评审描述,包含完整构建流程与规范说明。KG20C-QA基于KG20C构建,定义一组问答模板,将图谱三元组转化为自然语言问答对,支持基于图模型(如知识图谱嵌入)与文本模型(如大语言模型)的评估。我们在KG20C-QA上测试标准知识图谱嵌入方法,分析不同关系类型的性能表现,并提供可复现的评测协议。随论文发表,数据集将公开于https://github.com/tranhungnghiep/KG20C/,为学术问答、推理与知识驱动应用提供可复用、可扩展的研究资源。

原文摘要 · Abstract (English)

In this paper, we present KG20C and KG20C-QA, two curated datasets for advancing question answering (QA) research on scholarly data. KG20C is a high-quality scholarly knowledge graph constructed from the Microsoft Academic Graph through targeted selection of venues, quality-based filtering, and schema definition. Although KG20C has been available online in non-peer-reviewed sources such as GitHub repository, this paper provides the first formal, peer-reviewed description of the dataset, including clear documentation of its construction and specifications. KG20C-QA is built upon KG20C to support QA tasks on scholarly data. We define a set of QA templates that convert graph triples into natural language question--answer pairs, producing a benchmark that can be used both with graph-based models such as knowledge graph embeddings and with text-based models such as large language models. We benchmark standard knowledge graph embedding methods on KG20C-QA, analyze performance across relation types, and provide reproducible evaluation protocols. By officially releasing these datasets with thorough documentation, we aim to contribute a reusable, extensible resource for the research community, enabling future work in QA, reasoning, and knowledge-driven applications in the scholarly domain. The full datasets will be released at https://github.com/tranhungnghiep/KG20C/ upon paper publication.

知识图谱学术问答数据集推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。