arXiv:2608.30214cs.AI2026-08

用论文结构生成高难度科学推理数据,提升模型理解机制能力。

SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature

论文配图:SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature
图 1 · 摘自论文原文
  • 以论文的论点-证据-推导结构为单元,提炼紧凑推理骨架。
  • 构建包含23.4万条样本的Spark-234K数据集,涵盖四种科学推理视角。
  • 适合训练需要深度逻辑与证据推理的科学智能模型。

科学推理对开源模型仍是挑战,主要源于高质量科学推理数据的匮乏。现有数据集多聚焦事实回忆或公式化解题,缺乏对机制理解、证据支撑推理和假设检验的重视。为此,我们提出SPARK(Scientific Paper Abstracted Reasoning sKeleton),一个基于跨10个学科的Sci-Base大规模论文语料库的论文导向合成框架。SPARK不直接将论文转为问答对,而是将论文的论点-证据-推导结构视为推理合成的基本单元:(1)将每篇论文提炼为包含核心论点与支持证据的紧凑推理骨架,实现自洽问题生成;(2)从机制推理、假说证伪、定量推导和边界校准四个科学视角合成推理任务。最后通过一致性验证阶段剔除无支持或矛盾输出。基于此框架,我们构建了Spark-234K,一个难度与多样性显著高于现有资源的科学推理数据集。实验表明,Spark-234K在多个评测中持续优于现有数据集,并在使用更少训练样本时仍表现更强。

原文摘要 · Abstract (English)

Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism understanding, evidence-grounded reasoning, and hypothesis evaluation. To address this, we introduce SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, a large-scale corpus of research papers spanning 10 scientific disciplines. Instead of directly converting papers into question-answer pairs, SPARK treats the claim-evidence-derivation structure of a paper as the fundamental unit of reasoning synthesis. Specifically, SPARK (1) distills each paper into a compact reasoning skeleton capturing its central claims and supporting evidence, enabling self-contained question generation, and (2) synthesizes reasoning tasks from four scientific perspectives: mechanistic reasoning, hypothesis falsification, quantitative derivation, and boundary calibration. A final consistency verification stage further removes unsupported or contradictory outputs. Using this framework, we construct Spark-234K, a scientific reasoning dataset with substantially higher difficulty and diversity than existing resources. Experiments show that Spark-234K consistently outperforms existing scientific reasoning datasets while achieving stronger performance with significantly fewer training samples.

科学推理数据合成机制理解推理骨架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。