arXiv:2412.02788cs.CLcs.AI2024-12被引 5

构建首个融合文本与知识图谱的学术问答数据集

Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset

  • 用大模型生成10.5万条跨源问题,结合DBLP、SemOpenAlex和维基百科
  • 基于RAG的基线模型在测试集上达69.65%准确率
  • 适合研究多源信息融合的学者与开源系统开发者

现有学术问答方法通常仅依赖单一数据源,如纯文本或知识图谱(KG)。然而,学术信息常分布于异构来源,亟需能整合多源信息的问答系统。为此,我们提出 Hybrid-SQuAD(混合学术问答数据集),一个大规模数据集,用于支持融合文本与知识图谱事实的问答任务。该数据集包含10.5K个由大语言模型生成的问题-答案对,结合了DBLP与SemOpenAlex的知识图谱以及对应维基百科文本。此外,我们提出一种基于RAG的基线混合问答模型,在Hybrid-SQuAD测试集上取得69.65%的精确匹配分数。

原文摘要 · Abstract (English)

Existing Scholarly Question Answering (QA) methods typically target homogeneous data sources, relying solely on either text or Knowledge Graphs (KGs). However, scholarly information often spans heterogeneous sources, necessitating the development of QA systems that integrate information from multiple heterogeneous data sources. To address this challenge, we introduce Hybrid-SQuAD (Hybrid Scholarly Question Answering Dataset), a novel large-scale QA dataset designed to facilitate answering questions incorporating both text and KG facts. The dataset consists of 10.5K question-answer pairs generated by a large language model, leveraging the KGs DBLP and SemOpenAlex alongside corresponding text from Wikipedia. In addition, we propose a RAG-based baseline hybrid QA model, achieving an exact match score of 69.65 on the Hybrid-SQuAD test set.

学术问答知识图谱多源融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。