自动生成跨文献多跳科学问答数据集,助力科研推理能力评估
Automatic Inter-document Multi-hop Scientific QA Generation
- 用大模型提取单跳问题,结合语义对齐构建跨文献关联
- 从8211篇论文生成41万+单跳与1.37万个多跳问题
- 适合研究科学推理、检索增强问答的学者使用
现有自动科学问题生成研究多集中于单文档事实型问答,忽视了科学理解中至关重要的跨文档推理。我们提出AIM-SciQA,一个自动化生成多文档、多跳科学问答数据集的框架。该框架利用大语言模型(LLMs)进行机器阅读理解提取单跳问答,并基于嵌入的语义对齐构建跨文档关系,同时有选择性地利用引文信息。在8,211篇PubMed Central论文上应用,生成了411,409个单跳和13,672个多跳问题,形成IM-SciQA数据集。人工与自动验证表明其具有高事实一致性,实验结果证明该数据集能有效区分检索与问答阶段的推理能力,为检索增强型科学推理提供了真实且可解释的基准。我们进一步扩展该框架构建了CIM-SciQA,一种以引文引导的变体,在性能上接近理想设置,强化了数据集的有效性与通用性。
原文摘要 · Abstract (English)
Existing automatic scientific question generation studies mainly focus on single-document factoid QA, overlooking the inter-document reasoning crucial for scientific understanding. We present AIM-SciQA, an automated framework for generating multi-document, multi-hop scientific QA datasets. AIM-SciQA extracts single-hop QAs using large language models (LLMs) with machine reading comprehension and constructs cross-document relations based on embedding-based semantic alignment while selectively leveraging citation information. Applied to 8,211 PubMed Central papers, it produced 411,409 single-hop and 13,672 multi-hop QAs, forming the IM-SciQA dataset. Human and automatic validation confirmed high factual consistency, and experimental results demonstrate that IM-SciQA effectively differentiates reasoning capabilities across retrieval and QA stages, providing a realistic and interpretable benchmark for retrieval-augmented scientific reasoning. We further extend this framework to construct CIM-SciQA, a citation-guided variant achieving comparable performance to the Oracle setting, reinforcing the dataset's validity and generality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。