通过中间阶段训练提升模型化学推理能力,让强化学习有效起效。
MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models
- 设计中段科学训练(MiST)增强符号能力和隐含化学知识
- 使有机反应命名准确率从10.9%升至63.9%,无机材料生成从40.6%升至67.4%
- 适用于需可解释推理的化学建模任务,适合领域模型开发者
大语言模型可通过基于规则奖励的在线微调发展推理能力。然而近期研究发现,强化学习仅在基础模型对正确答案已赋予非忽略概率时才有效——我们称此为‘隐含可解性’。本文探究化学推理能力的出现及其对化学领域的意义。识别出基于强化学习的化学推理两大必要条件:1)符号能力;2)隐含化学知识。提出中段科学训练(MiST):一系列中段训练技术,包括结合SMILES/CIF感知预处理的数据混合、29亿token的持续预训练,以及10亿token的监督微调。这些步骤将3B和7B模型的隐含可解性得分提升最多1.8倍,并使强化学习在有机反应命名任务中将准确率从10.9%提升至63.9%,在无机材料生成任务中从40.6%提升至67.4%。其他挑战性化学任务也获得类似结果,且生成可解释的推理轨迹。研究明确了化学推理训练的关键前提,凸显中段训练在激发推理能力中的关键作用。
原文摘要 · Abstract (English)
Large Language Models can develop reasoning capabilities through online fine-tuning with rule-based rewards. However, recent studies reveal a critical constraint: reinforcement learning succeeds only when the base model already assigns non-negligible probability to correct answers -- a property we term 'latent solvability'. This work investigates the emergence of chemical reasoning capabilities and what these prerequisites mean for chemistry. We identify two necessary conditions for RL-based chemical reasoning: 1) Symbolic competence, and 2) Latent chemical knowledge. We propose mid-stage scientific training (MiST): a set of mid-stage training techniques to satisfy these, including data-mixing with SMILES/CIF-aware pre-processing, continued pre-training on 2.9B tokens, and supervised fine-tuning on 1B tokens. These steps raise the latent-solvability score on 3B and 7B models by up to 1.8x, and enable RL to lift top-1 accuracy from 10.9 to 63.9% on organic reaction naming, and from 40.6 to 67.4% on inorganic material generation. Similar results are observed for other challenging chemical tasks, while producing interpretable reasoning traces. Our results define clear prerequisites for chemical reasoning training and highlight the broader role of mid-stage training in unlocking reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。