构建125万条科学推理数据集,提升AI科研能力
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
- 整合1.25万条高质量科学问题,覆盖7大学科
- 在15个基准上表现超越现有开源数据集
- 适合大模型科学推理训练,支持研究复现
科学推理对培养AI科学家和辅助人类科研至关重要。然而,开源社区长期聚焦数学与编程,忽视科学领域,主要因缺乏开放、大规模、高质量且可验证的科学推理数据集。为此,我们提出TextbookReasoning,一个包含12,000所大学教材提取的65万条真实答案的开源数据集,覆盖7个科学领域。进一步构建MegaScience,整合125万条高质量开源数据,通过系统消融实验优化各数据集的选取策略。同时建立涵盖15个基准、多种题型的评估体系,结合精准答案提取方法确保评测准确。实验表明,该数据集在性能和训练效率上优于现有开源科学数据集,响应更简洁。在Llama3.1、Qwen2.5及Qwen3系列基模型上微调后,显著优于对应官方指令模型。且对更大更强模型更具效果,体现科学微调的缩放优势。已公开数据处理流程、评估系统、数据集及七个训练模型,推动科学推理研究发展。
原文摘要 · Abstract (English)
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning datasets. To bridge this gap, we first present TextbookReasoning, an open dataset featuring truthful reference answers extracted from 12k university-level scientific textbooks, comprising 650k reasoning questions spanning 7 scientific disciplines. We further introduce MegaScience, a large-scale mixture of high-quality open-source datasets totaling 1.25 million instances, developed through systematic ablation studies that evaluate various data selection methodologies to identify the optimal subset for each publicly available scientific dataset. Meanwhile, we build a comprehensive evaluation system covering diverse subjects and question types across 15 benchmarks, incorporating comprehensive answer extraction strategies to ensure accurate evaluation metrics. Our experiments demonstrate that our datasets achieve superior performance and training efficiency with more concise response lengths compared to existing open-source scientific datasets. Furthermore, we train Llama3.1, Qwen2.5, and Qwen3 series base models on MegaScience, which significantly outperform the corresponding official instruct models in average performance. In addition, MegaScience exhibits greater effectiveness for larger and stronger models, suggesting a scaling benefit for scientific tuning. We release our data curation pipeline, evaluation system, datasets, and seven trained models to the community to advance scientific reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。