arXiv:2603.29723cs.AI2026-03被引 1

让化学合成路径生成像解题一样一步步推理,直接端到端优化整体策略。

Reinforced Reasoning for End-to-End Retrosynthetic Planning

  • 将逆合成看作链式思考过程,用逐步推理替代传统分步预测+搜索的割裂方式。
  • 在RetroBench上达到新SOTA,长路径规划鲁棒性显著优于混合基线方法。
  • 通过可验证奖励的强化学习训练,让每一步生成都服务于最终合成可行性。

逆合成规划是有机化学中的基础任务,但由于其组合复杂性仍具挑战。传统方法通常依赖将单步预测与外部搜索启发式结合的混合框架,不可避免地破坏了局部分子转化与全局规划目标之间的逻辑连贯性。为弥合这一差距,并将复杂的策略远见直接嵌入模型的化学推理中,我们提出ReTriP,一种端到端生成框架,将逆合成重新建模为直接的链式思考推理任务。我们构建了路径一致的分子表示,并采用渐进式训练流程,从推理蒸馏过渡到带有可验证奖励的强化学习,有效对齐逐步生成与实际路线效用。在RetroBench上的实证评估显示,ReTriP达到当前最优性能,在长程规划中展现出更强的鲁棒性,优于混合基线方法。

原文摘要 · Abstract (English)

Retrosynthetic planning is a fundamental task in organic chemistry, yet remains challenging due to its combinatorial complexity. To address this, conventional approaches typically rely on hybrid frameworks that combine single-step predictions with external search heuristics, inevitably fracturing the logical coherence between local molecular transformations and global planning objectives. To bridge this gap and embed sophisticated strategic foresight directly into the model's chemical reasoning, we introduce ReTriP, an end-to-end generative framework that reformulates retrosynthesis as a direct Chain-of-Thought reasoning task. We establish a path-coherent molecular representation and employ a progressive training curriculum that transitions from reasoning distillation to reinforcement learning with verifiable rewards, effectively aligning stepwise generation with practical route utility. Empirical evaluation on RetroBench demonstrates that ReTriP achieves state-of-the-art performance, exhibiting superior robustness in long-horizon planning compared to hybrid baselines.

逆合成链式推理强化学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。