arXiv:2509.21320cs.CL2025-09被引 10

跨学科科学推理大模型,能理解并生成各类科学内容。

SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines

  • 用2060亿文本训练,结合指令微调与强化学习,实现科学思维链推理。
  • 支持103项任务,跨领域泛化能力优于专用系统,生成更准确。
  • 适合科研人员、教育者,用于科学内容生成与知识提取。

我们提出一种科学推理基础模型,将自然语言与异构科学表示对齐。模型在包含2060亿个标记的语料库上预训练,涵盖科学文本、纯序列及序列-文本对,随后通过4000万条指令进行监督微调,并采用渐进式冷启动自举方法激发长链思维过程,再结合任务特定奖励塑造的强化学习,赋予模型刻意的科学推理能力。该模型支持五大能力族,覆盖最多达103项任务,涵盖工作流中的:(i) 文本与科学格式间的忠实转换,(ii) 文本/知识提取,(iii) 属性预测,(iv) 属性分类,(v) 无条件与有条件序列生成与设计。相较于专用系统,本方法扩大了指令覆盖范围,提升了跨域泛化能力,并增强输出保真度。我们详述数据构建与训练流程,并证明跨学科学习可提升迁移性能与下游可靠性。模型、指令微调数据集与评估代码已开源至 https://huggingface.co/SciReason 及 https://github.com/open-sciencelab/SciReason。

原文摘要 · Abstract (English)

We present a scientific reasoning foundation model that aligns natural language with heterogeneous scientific representations. The model is pretrained on a 206B-token corpus spanning scientific text, pure sequences, and sequence-text pairs, then aligned via SFT on 40M instructions, annealed cold-start bootstrapping to elicit long-form chain-of-thought, and reinforcement learning with task-specific reward shaping, which instills deliberate scientific reasoning. It supports four capability families, covering up to 103 tasks across workflows: (i) faithful translation between text and scientific formats, (ii) text/knowledge extraction, (iii) property prediction, (iv) property classification, (v) unconditional and conditional sequence generation and design. Compared with specialist systems, our approach broadens instruction coverage, improves cross-domain generalization, and enhances fidelity. We detail data curation and training and show that cross-discipline learning strengthens transfer and downstream reliability. The model, instruct tuning datasets and the evaluation code are open-sourced at https://huggingface.co/SciReason and https://github.com/open-sciencelab/SciReason.

科学推理多模态大模型开放数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。