提升分子大模型推理能力,让化学分析更准确可解释。
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- 两阶段训练:先用增强思维链数据微调,再用任务自适应奖励优化。
- 在分子生成与描述任务中显著优于多个基线模型。
- 设计更可解释,减少化学结构与文本幻觉问题,适合药物研发者使用。
大语言模型在多个领域表现优异,但在分子推理方面仍不充分。现有方法多依赖通用提示,缺乏领域特异性化学语义;或通过微调,但存在可解释性差、推理深度不足,常引发结构与文本幻觉。为此,我们提出 MolReasoner,一种两阶段框架,使模型从记忆转向高保真化学推理。第一阶段(Mol-SFT)利用知识增强的思维链(CoT)数据建立坚实基础;第二阶段(Mol-RL)通过新颖的任务自适应奖励机制优化推理,缓解幻觉。大量评估显示,MolReasoner 在分子生成和描述任务中显著超越多种强基线。进一步分析表明其设计协同性强,输出更具可解释性。本工作为实现高保真分子推理提供了系统且有效的新路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive performance across various domains, but their ability to perform molecular reasoning remains underexplored. Existing methods mostly rely on general-purpose prompting, which lacks domain-specific molecular semantics, or fine-tuning, which faces challenges in interpretability and reasoning depth, often leading to structural and textual hallucinations. To address these issues, we introduce MolReasoner, a two-stage framework that transitions LLMs from memorization to high-fidelity chemical reasoning. In the Mol-SFT stage, knowledge-enhanced Chain-of-Thought (CoT) data provides a strong foundation, while the Mol-RL stage refines reasoning using a novel, task-adaptive reward system to mitigate hallucinations. Extensive evaluations demonstrate that MolReasoner significantly outperforms a wide range of strong baselines in both molecule generation and captioning tasks. Further analyses highlight the framework's synergistic design and its ability to produce more interpretable outputs. Our work presents a principled and effective new approach for advancing high-fidelity molecular reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。