arXiv:2604.07669cs.LGcs.AI2026-04被引 1

用化学可行反应模板约束生成,让药物分子优化更高效且可合成。

Reinforcement Learning with LLM-Guided Action Spaces for Synthesizable Lead Optimization

论文配图:Reinforcement Learning with LLM-Guided Action Spaces for Synthesizable Lead Optimization
图 1 · 摘自论文原文
  • 基于验证反应模板构建合成约束动作空间,确保每步修改都可实现。
  • 在14项任务中平均排名前10得分为0.571,9项任务样本效率最优。
  • 结合工具增强的LLM与强化学习,兼顾化学合理性与优化性能。

药物先导化合物优化需在提升治疗特性的同时保证分子结构修改具有可行合成路径。现有方法或忽视可合成性,或依赖昂贵的反应网络枚举,直接使用大语言模型生成常出现化学无效结构。我们提出MolReAct框架,将优化问题建模为基于合成约束动作空间的马尔可夫决策过程,该空间由经验证的反应模板定义。一个工具增强的LLM代理作为动态反应环境,调用专业化学分析工具识别反应位点和官能团,从匹配模板中提出少量化学合理的转化策略。通过组相对策略优化(GRPO)训练专用策略模型,在多步轨迹中最大化长期预言奖励,采用基于SMILES的缓存机制使端到端优化时间减少约43%。在来自治疗数据共同体的13项属性优化任务及一项基于结构的对接任务中,MolReAct平均Top-10得分达0.571,所有基线中最高,在14项任务中有13项排名第一或第二,9项任务中样本效率最佳。通过每一步均基于验证模板,所生成分子不仅性质更优,且每条路径均有明确的模板支撑合成路线。

原文摘要 · Abstract (English)

Lead optimization in drug discovery requires improving therapeutic properties while ensuring that molecular modifications correspond to feasible synthetic routes. Existing approaches either prioritize property scores without enforcing synthesizability, or rely on expensive enumeration over large reaction networks, while direct application of Large Language Models (LLMs) to molecular generation frequently produces chemically invalid structures. We introduce MolReAct, a framework that formulates lead optimization as a Markov Decision Process over a synthesis-constrained action space defined by validated reaction templates. A tool-augmented LLM agent serves as a dynamic reaction environment, invoking specialized chemical analysis tools to identify reactive sites and functional groups and proposing a compact set of chemically grounded transformations from matched templates. A dedicated policy model trained via Group Relative Policy Optimization (GRPO) selects among these constrained actions to maximize long-term oracle reward across multi-step trajectories, with a SMILES-based caching mechanism reducing end-to-end optimization time by approximately 43%. Across 13 property optimization tasks from the Therapeutic Data Commons and one structure-based docking task, MolReAct achieves an average Top-10 score of 0.571, the highest among all baselines, ranking first or second on 13 of 14 tasks and attaining the best sample efficiency on 9 of 14 tasks. By grounding every optimization step in validated reaction templates, MolReAct produces molecules that are not only property-improved but each accompanied by an explicit template-grounded synthetic pathway.

药物发现强化学习生成化学可合成性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。