arXiv:2501.15587cs.CLcs.AI2025-01被引 31

构建11.6万条高质量高教科学题解数据集,助力大模型科学推理能力提升

SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain

  • 基于通用化流水线,从异构来源自动提取题解对
  • 包含116,756组题解,经严格筛选确保科学严谨性与教学层级
  • 开源数据集与代码,适合科研人员开展科学推理研究

大型语言模型(LLMs)在数学与科学推理方面取得突破,尤其o1模型展现了强大能力,凸显高质量训练数据对推动STEM领域模型性能的关键作用。尽管数学领域已有丰富数据资源,高等教育科学领域长期缺乏相应优质数据集。为此,我们提出SCP-116K,一个包含116,756组高质量问题-解答对的大规模数据集,通过简化且高度可泛化的自动化提取流程,从异构来源中获取。该方法结合严格过滤机制,确保内容的科学严谨性与教育适切性,同时具备良好的可扩展性与跨领域迁移潜力。我们公开发布数据集及提取流水线,旨在促进科学推理研究,支持新模型的全面评估,并降低其他研究团队复现o1等先进模型成果的门槛。我们认为SCP-116K将成为推动高层次科学推理任务的重要资源,加速大模型在科学领域的创新。数据集与代码已开源:https://github.com/AQA6666/SCP-116K-open。

原文摘要 · Abstract (English)

Recent breakthroughs in large language models (LLMs) exemplified by the impressive mathematical and scientific reasoning capabilities of the o1 model have spotlighted the critical importance of high-quality training data in advancing LLM performance across STEM disciplines. While the mathematics community has benefited from a growing body of curated datasets, the scientific domain at the higher education level has long suffered from a scarcity of comparable resources. To address this gap, we present SCP-116K, a new large-scale dataset of 116,756 high-quality problem-solution pairs, automatically extracted from heterogeneous sources using a streamlined and highly generalizable pipeline. Our approach involves stringent filtering to ensure the scientific rigor and educational level of the extracted materials, while maintaining adaptability for future expansions or domain transfers. By openly releasing both the dataset and the extraction pipeline, we seek to foster research on scientific reasoning, enable comprehensive performance evaluations of new LLMs, and lower the barrier to replicating the successes of advanced models like o1 in the broader science community. We believe SCP-116K will serve as a critical resource, catalyzing progress in high-level scientific reasoning tasks and promoting further innovations in LLM development. The dataset and code are publicly available at https://github.com/AQA6666/SCP-116K-open.

科学推理数据集大模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。