arXiv:2607.04688cs.LGcs.AI2026-07

提出化学可解释的逆合成评估框架,对比大模型与专用模型性能

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

论文配图:URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
图 1 · 摘自论文原文
  • 构建兼顾形式与化学合理性的逆合成评估体系
  • 大模型擅长战略规划但难可靠完成合成路径
  • 针对真实药物研发任务测试新基准

旨在为靶分子寻找反应路径的合成规划是药物发现中最重要的挑战之一。近期已有专门的深度学习逆合成系统和通用大语言模型(LLMs)取得进展,但由于缺乏灵活且化学可解释的评估协议,客观比较仍很困难。本研究提出了URSA(Utilitarian RetroSynthesis Assessment)评估框架,不仅从收敛到商业化起始原料的形式角度,还从化学合理性角度评估合成路线,模拟专家评价方式。在一组新颖、多样且合成路线未公开的靶分子上,对传统端到端逆合成方案和大语言模型进行了全面评估,这些目标代表了日常药物设计中的真实任务。结果表明,尽管大语言模型可在高层次战略规划中提供支持,但在可靠解决合成规划任务方面,目前仍不如专用逆合成模型。

原文摘要 · Abstract (English)

Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Recent progress has produced both specialized deep-learning retrosynthesis systems and general-purpose large language models, but objective comparison remains difficult due to the lack of flexible, chemically interpretable benchmarking protocols. In the current study, we are introducing the URSA (Utilitarian RetroSynthesis Assessment) evaluation framework that provides the opportunity to benchmark the synthetic routes not only from a formal perspective, such as convergence to commercially available starting materials, but also from a chemical plausibility perspective, mimicking the way expert chemists evaluate the reactions and routes. The study covers a comprehensive evaluation of both conventional end-to-end retrosynthesis solutions and LLMs for the synthesis planning task on a set of novel, diverse target molecules with undisclosed synthetic routes, which represent realistic tasks in the daily drug design routine. We find that while LLMs can support high-level strategic planning, they currently underperform specialized retrosynthesis models in reliably solving synthesis planning tasks.

逆合成评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。