arXiv:2603.14257cs.CL2026-03中稿 · the 2026 Internati…被引 2

自动生成跨文献多跳科学问答数据集,助力科研推理能力评估

Automatic Inter-document Multi-hop Scientific QA Generation

  • 用大模型提取单跳问题,结合语义对齐构建跨文献关联
  • 从8211篇论文生成41万+单跳与1.37万个多跳问题
  • 适合研究科学推理、检索增强问答的学者使用

现有自动科学问题生成研究多集中于单文档事实型问答,忽视了科学理解中至关重要的跨文档推理。我们提出AIM-SciQA,一个自动化生成多文档、多跳科学问答数据集的框架。该框架利用大语言模型(LLMs)进行机器阅读理解提取单跳问答,并基于嵌入的语义对齐构建跨文档关系,同时有选择性地利用引文信息。在8,211篇PubMed Central论文上应用,生成了411,409个单跳和13,672个多跳问题,形成IM-SciQA数据集。人工与自动验证表明其具有高事实一致性,实验结果证明该数据集能有效区分检索与问答阶段的推理能力,为检索增强型科学推理提供了真实且可解释的基准。我们进一步扩展该框架构建了CIM-SciQA,一种以引文引导的变体,在性能上接近理想设置,强化了数据集的有效性与通用性。

原文摘要 · Abstract (English)

Existing automatic scientific question generation studies mainly focus on single-document factoid QA, overlooking the inter-document reasoning crucial for scientific understanding. We present AIM-SciQA, an automated framework for generating multi-document, multi-hop scientific QA datasets. AIM-SciQA extracts single-hop QAs using large language models (LLMs) with machine reading comprehension and constructs cross-document relations based on embedding-based semantic alignment while selectively leveraging citation information. Applied to 8,211 PubMed Central papers, it produced 411,409 single-hop and 13,672 multi-hop QAs, forming the IM-SciQA dataset. Human and automatic validation confirmed high factual consistency, and experimental results demonstrate that IM-SciQA effectively differentiates reasoning capabilities across retrieval and QA stages, providing a realistic and interpretable benchmark for retrieval-augmented scientific reasoning. We further extend this framework to construct CIM-SciQA, a citation-guided variant achieving comparable performance to the Oracle setting, reinforcing the dataset's validity and generality.

科学问答多跳推理数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。