arXiv:2607.12771cs.LGcs.CE2026-07

用大模型学习化学反应机理,提升推理能力。

Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models

论文配图:Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models
图 1 · 摘自论文原文
  • 构建大规模机理推理数据集,支持逐步推理解析。
  • 在FukuyamaBench上达8.3%路径匹配率,优于专用模型。
  • 适合需要深度化学推理的科研与教育场景。

反应机理是解释化学转化的逐步基元反应序列。学习机理逻辑对提升大语言模型(LLM)的基础化学智能至关重要。反应机理的逐步推导天然契合推理型大模型的思维范式。然而,现有化学LLM多聚焦于粗粒度名称反应的产物预测与逆合成,常导致物理不一致和幻觉。相比之下,专用于机理推断的小规模生成模型普遍泛化能力受限。为此,我们构建了新的大规模反应机理推理数据集,并建立了源自《福山高级有机反应机理》一书的FukuyamaBench基准,用于严格评估模型在分层机理推理上的表现。微调后的Qwen3-30B-A3B在FukuyamaBench Set~A上达到8.3%的精确路径匹配率,超过专用模型FlowER的5.1%,表明机理感知训练显著提升了语言模型的化学推理能力。

原文摘要 · Abstract (English)

Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations. Learning the mechanism logic is therefore essential for enhancing the fundamental chemical intelligence of large language models (LLMs). The stepwise deduction of reaction mechanism aligns naturally with the reasoning paradigms of reasoning LLMs. However, current chemical LLMs primarily emphasize coarse-grained name reactions for product prediction and retrosynthesis, often leading to physical inconsistencies and hallucinations. In contrast, specialized small-scale generative models for mechanism inference typically suffer from restricted generalization capacity across diverse chemical spaces. To overcome these limitations, we built a novel, large-scale reasoning dataset of reaction mechanisms. Furthermore, we established the FukuyamaBench, a difficult benchmark derived from Fukuyama's Advanced Organic Reaction Mechanism book, to rigorously evaluate model performance on hierarchical mechanism reasoning. Our fine-tuned Qwen3-30B-A3B achieves 8.3% exact pathway match on FukuyamaBench Set~A, surpassing the specialized FlowER model (5.1%), demonstrating that mechanism-aware training substantially enhances chemical reasoning in language models.

化学推理大模型机理学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。