arXiv:2608.30702cs.CLcs.AI2026-08

用强化学习选分子路径,提升生物合成逆向设计效率

An Agentic Retrobiosynthesis Framework with Learned Frontier Selection

论文配图:An Agentic Retrobiosynthesis Framework with Learned Frontier Selection
图 1 · 摘自论文原文
  • 用大模型选择下一步要分解的分子,不改变反应规则
  • 在有限计算预算下,准确率最高达88%,超越传统方法
  • 适合做代谢工程和药物分子设计的科研人员参考

大型语言模型被用于多步逆向生物合成任务,但其搜索策略的作用是否独立于反应模型仍存疑。本研究在生物合成场景中通过基于规则的逆向生物合成框架进行验证:一个确定性的生化引擎生成所有方法一致的反应路径,仅由策略决定下一步扩展哪个前体分子。使用提示工程和LoRA微调的Qwen2.5-7B模型采用严格的选择接口。微调后的策略在LASER数据集上10步内达到65±1%的解决率,优于MCTS的59%;200步时达到78±1%,高于LASER的75%、RetroPath RL黄金基准的88±3%,以及BioNavi-NP基准的63±2%(对比45%)。微调始终优于直接提示。结果表明,路线监督的前体选择可在不改变生化生成机制的前提下提升预算受限的搜索性能,但表现仍依赖前体构建与反应排序质量。

原文摘要 · Abstract (English)

Large language models are increasingly used as agents for multistep retrosynthesis, raising the question of how much their search policy contributes independently of the underlying reaction model. We investigate this question in a biological setting through rule-based retrobiosynthesis: a deterministic biochemical engine generates the same validated transitions for every method, searching for routes that terminate in metabolites available to an \emph{Escherichia coli} chassis, while the policy only selects which frontier molecule to expand next. Prompted and LoRA-tuned Qwen2.5-7B policies use a strict choice-only interface. The fine-tuned policy reaches $65\pm1$\% solve rate at 10 expansions on LASER versus 59\% for MCTS, and at 200 expansions reaches $78\pm1$\% versus 75\% on LASER, $88\pm3$\% versus 80\% on the RetroPath RL Golden benchmark, and $63\pm2$\% versus 45\% on the BioNavi-NP benchmark. Fine-tuning also consistently outperforms direct prompting. These results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.

逆向合成大模型应用代谢工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。