arXiv:2507.07328cs.LGcs.AI2025-07被引 1

用推理增强模型解决化学生成中的看似合理实则错误问题

Bridging the Plausibility-Validity Gap by Fine-Tuning a Reasoning-Enhanced LLM for Chemical Synthesis and Discovery

  • 用推理模型+低秩微调,结合分子属性与反应数据训练
  • 化学有效率97.4%,合成可行性达74.4%,格式正确率96.3%
  • 比MolT5更准,接近复杂系统表现,过程透明高效

大型语言模型常生成看似科学合理却违背基本原理的输出,我们称之为“可信性-有效性鸿沟”。这一问题在化学领域尤为严重,表面正确的分子结构、反应机理和合成路径可能隐藏深层错误。本文提出一种系统方法:采用以推理为中心的模型架构(Magistral Small),在涵盖分子性质与化学转化的双域数据集上进行低秩适应微调。评估显示,微调后系统达到96.3%的格式一致率、97.4%的化学有效性及74.4%的合成可行性。相比专用翻译模型MolT5(97.4% vs 77.2%有效性),本方法表现更优;且在专家评分(9.0/10)上接近复杂工具增强系统ChemCrow(9.24/10),同时具备更透明、高效的特性。结果揭示学习层次:语法正确性先于化学理解,后者又先于合成规划能力。该工作建立可复现的通用模型转化框架,指出立体化学精度、知识时效性与计算可及性为未来关键挑战。

原文摘要 · Abstract (English)

Large Language Models frequently generate outputs that appear scientifically reasonable yet violate fundamental principles--a phenomenon we characterize as the "plausibility-validity gap." This challenge proves especially acute in chemistry, where superficial correctness masks deeper errors in molecular structure, reaction mechanisms, and synthetic pathways. We present a systematic approach combining a reasoning-centric model architecture (Magistral Small) with Low-Rank Adaptation fine-tuning on a dual-domain dataset covering molecular properties and chemical transformations. Evaluation reveals substantial improvements: the fine-tuned system achieves 96.3% format adherence, 97.4% chemical validity, and 74.4% synthesis feasibility. Comparative analysis shows our approach outperforms specialized translation models like MolT5 (97.4% vs 77.2% validity) while achieving performance comparable to complex tool-augmented systems like ChemCrow (9.0/10 vs 9.24/10 expert rating) through a more transparent, efficient methodology. Results demonstrate a learning hierarchy where syntactic correctness develops before chemical understanding, which precedes synthetic planning capability. This work establishes a reproducible framework for transforming generalist language models into dependable scientific tools while identifying critical areas including stereochemical precision, knowledge currency, and computational accessibility as key challenges for future advancement.

化学生成大模型微调推理增强有效性提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。