用强化学习让大模型学会化学推理,自动生成可解释的合成路径。
RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
- 基于强化学习训练大模型,实现可解释的化学逆合成推理
- 在USPTO-50K上达到66.1%准确率,优于现有方法
- 能生成真实药物与材料的多步合成路径,适合科研与工业应用
逆合成规划是有机合成与药物发现的核心。现有AI方法多依赖模式匹配,缺乏可迁移的化学推理能力,限制了泛化性与可解释性。本文提出RetroDFM-R,一种基于强化学习的推理驱动型大语言模型,用于化学逆合成预测。该模型不仅提升准确性,还提供透明、分步的推理过程。在USPTO-50K基准测试中,无增强情况下准确率达60.4%,全推理设置下达66.1%,超越当前最优基线。双盲专家评估进一步验证其路径的化学合理性与实用性。我们还证明RetroDFM-R可重构真实药物及自组装单层材料的复杂多步合成路线。通过显式表达推理过程,该模型克服了信任障碍,支持自动化逆合成规划的实际部署。
原文摘要 · Abstract (English)
Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existing AI methods often rely on pattern matching rather than transferable chemical reasoning, limiting both generalizability and interpretability. Here we introduce RetroDFM-R, a reasoning-driven large language model (LLM) for chemical retrosynthesis. Leveraging large-scale reinforcement learning, RetroDFM-R moves beyond black-box prediction by coupling improved accuracy with transparent, step-by-step rationale. On the USPTO-50K benchmark, RetroDFM-R achieves 60.4% accuracy without augmentation and 66.1% with the full inference setup, outperforming previous state-of-the-art baselines. Beyond standard metrics, double-blind expert evaluation further supports the chemical plausibility and practical utility of its proposed pathways. We also demonstrate that RetroDFM-R can reconstruct complex, multistep synthetic routes for real-world pharmaceuticals and self-assembled monolayer materials. By making its reasoning explicit and human-interpretable, RetroDFM-R addresses a key barrier to trust and supports practical deployment in automated retrosynthetic planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。