arXiv:2602.08321cs.CL2026-02被引 2

构建科学推理数据集与训练流程,显著提升大模型解题能力

Improving Data and Reward Design for Scientific Reasoning in Large Language Models

  • 构建100万题的跨学科科学数据集,支持可验证与开放答案区分
  • 在GPQA-diamond上达63.2分,超越o1-mini和GPT-4o等基线
  • 通过动态难度课程与评分标准引导强化学习,适合科研与教育场景

解决开放式科学问题对大语言模型仍是挑战,主要源于监督信号和评估机制不可靠。瓶颈在于科学后训练阶段的数据构建与奖励设计。本文提出大规模系统化数据处理流程,将异构开源科学数据转化为Dr. SCI数据集,包含8个STEM领域共100万道题目,具备明确可验证/开放答案划分、可扩展难度标注及细粒度评分标准,可操作化评估开放答案。基于此数据集,提出Dr. SCI后训练流程,重构标准SFT→RL范式,包含三部分:(i)探索扩展型SFT,提升模型推理模式覆盖;(ii)动态难度课程,随模型能力自适应调整数据;(iii)SciRubric引导的强化学习,通过评分标准实现开放问题上的稳定奖励。使用该流程训练的Qwen3-4B-Base在GPQA-diamond上达到63.2分,在GPQA-general上达32.4分,持续优于o1-mini和GPT-4o等强基线,尤其在开放题中表现显著提升。

原文摘要 · Abstract (English)

Solving open-ended science questions remains challenging for large language models, particularly due to inherently unreliable supervision and evaluation. The bottleneck lies in the data construction and reward design for scientific post-training. We develop a large-scale, systematic data processing pipeline that transforms heterogeneous open-source science data into Dr. SCI dataset, which comprises of 1M questions across eight STEM subjects, with explicit verifiable/open-ended splits, scalable difficulty annotation, and fine-grained rubrics that operationalize evaluation for open-ended answers. Building on this dataset, we propose the Dr. SCI post-training pipeline, which redesigns the standard SFT -> RL workflow through three components: (i) Exploration-Expanding SFT, which broadens the model's reasoning pattern coverage prior to RL; (ii) Dynamic Difficulty Curriculum, which adapts training data to the model's evolving scientific capability; and (iii) SciRubric-Guided RL, which enables stable reinforcement learning on open-ended scientific questions via rubric-based evaluation with explicit answer correctness. Qwen3-4B-Base trained using Dr. SCI pipeline achieves 63.2 on GPQA-diamond and 32.4 on GPQA-general, consistently improves over strong post-trained baselines such as o1-mini and GPT-4o, demonstrating substantial gains in scientific reasoning, especially in open-ended settings.

科学推理大模型训练数据构建强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。