把逆合成规划看作最差路径优化,提升路线可靠性。
Retrosynthesis Planning via Worst-path Policy Optimisation in Tree-structured MDPs
- 将逆合成建模为树状MDP中的最差路径优化问题
- 在Retro*-190上100%解决目标分子,路线缩短4.9%
- 仅用10%数据达顶尖性能,适合追求鲁棒性的合成设计
逆合成规划旨在将目标分子分解为可获得的构建单元,形成合成树,其中内部节点代表中间化合物,叶节点理想对应可购买试剂。然而,若任一叶节点无效,整条合成路线即失效,使规划过程对“最弱环节”极为敏感。现有方法通常优化各分支的平均表现,未能考虑这种最坏情况下的脆弱性。本文将逆合成重构为树状马尔可夫决策过程(Tree-structured MDPs)中的最差路径优化问题。我们证明该形式存在唯一最优解,并具备单调改进保证。基于此,提出交互式逆合成规划(InterRetro),通过与树状MDP交互,学习最差路径的值函数,并利用自我模仿机制,优先强化过去具有高估计优势的决策。实验表明,InterRetro在Retro*-190基准上实现100%目标解决率,合成路线平均缩短4.9%,且仅使用10%训练数据即取得优异性能。
原文摘要 · Abstract (English)
Retrosynthesis planning aims to decompose target molecules into available building blocks, forming a synthetic tree where each internal node represents an intermediate compound and each leaf ideally corresponds to a purchasable reactant. However, this tree becomes invalid if any leaf node is not a valid building block, making the planning process vulnerable to the "weakest link" in the synthetic route. Existing methods often optimise for average performance across branches, failing to account for this worst-case sensitivity. In this paper, we reframe retrosynthesis as a worst-path optimisation problem within tree-structured Markov Decision Processes (MDPs). We prove that this formulation admits a unique optimal solution and provides monotonic improvement guarantees. Building on this insight, we introduce Interactive Retrosynthesis Planning (InterRetro), a method that interacts with the tree MDP, learns a value function for worst-path outcomes, and improves its policy through self-imitation, preferentially reinforcing past decisions with high estimated advantage. Empirically, InterRetro achieves state-of-the-art results - solving 100% of targets on the Retro*-190 benchmark, shortening synthetic routes by 4.9%, and achieving promising performance using only 10% of the training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。