arXiv:2506.11117cs.CLcs.AI2025-06KDD被引 10

生成6.1万条真实科研场景的问答数据,提升科学检索工具的实用能力。

ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research

  • 基于论文内容自动提取数据信息,构建更真实的科研知识表示。
  • 用认知分类框架生成6.1万条高质量复杂问题,覆盖真实研究需求。
  • 通过大模型困惑度变化自动筛选答案,贴近人类对答案有效性的判断。

科研人员需要深入理解数据集信息以评估和开发理论与方法,但这些信息需求通常隐含在具体研究任务中,而非直接体现在搜索查询中。现有科学检索与问答数据集多针对简单问题,难以反映真实科研中的复杂询问分布。为此,我们提出ScIRGen——一个面向科学问答与检索的数据集生成框架,用于构建大规模、真实感强的科学检索增强生成(RAG)数据集。技术上,设计了基于论文的数据集信息提取方法,增强数据表示;提出基于认知分类的问题生成框架,确保合成问题质量;并设计一种基于大模型困惑度变化的自动答案过滤方法,其效果与人类对答案有效性的判断高度一致。最终生成包含6.1万条样本的ScIRGen-Geo数据集。在该数据集上对主流方法进行基准测试,发现当前模型在处理复杂问题时仍存在推理短板。本工作推动了更复杂工具的发展,以满足科学界的信息需求。

原文摘要 · Abstract (English)

Scientific researchers need intensive information about datasets to effectively evaluate and develop theories and methodologies. The information needs regarding datasets are implicitly embedded in particular research tasks, rather than explicitly expressed in search queries. However, existing scientific retrieval and question-answering (QA) datasets typically address straightforward questions, which do not align with the distribution of real-world research inquiries. To bridge this gap, we developed ScIRGen, a dataset generation framework for scientific QA \& retrieval that more accurately reflects the information needs of professional science researchers, and uses it to create a large-scale scientific retrieval-augmented generation (RAG) dataset with realistic queries, datasets and papers. Technically, we designed a dataset-oriented information extraction method that leverages academic papers to augment the dataset representation. We then proposed a question generation framework by employing cognitive taxonomy to ensure the quality of synthesized questions. We also design a method to automatically filter synthetic answers based on the perplexity shift of LLMs, which is highly aligned with human judgment of answers' validity. Collectively, these methodologies culminated in the creation of the 61k QA dataset, ScIRGen-Geo. We benchmarked representative methods on the ScIRGen-Geo dataset for their question-answering and retrieval capabilities, finding out that current methods still suffer from reasoning from complex questions. This work advances the development of more sophisticated tools to support the intricate information needs of the scientific community.

科学问答数据集生成RAG大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。