arXiv:2602.03554cs.LGcs.AI2026-02被引 3

提出新评估框架,让大模型合成规划更贴近真实化学实践。

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

  • 用化学合理性指标替代单一答案匹配,更符合实际合成思路。
  • 构建百万级验证数据集,训练出性能优于基线的模型。
  • 适合关注药物设计与大模型评估的科研人员参考。

近期进展推动大语言模型(LLMs)在药物发现中的应用,包括合成路径规划。然而,对逆合成性能的客观评估仍有限。现有基准和指标通常依赖已发表的合成方案及基于单个标准答案的Top-K准确率,未能反映真实合成规划的开放性。本文提出一种新的单步逆合成评估框架,采用ChemCensor这一新型化学合理性度量,评估通用型与化学专用型LLMs。通过强调合理性而非精确匹配,该方法更契合人类合成规划实践。我们还引入CREED数据集,包含数百万经ChemCensor验证的反应记录,用于训练模型,并在该基准下实现对现有LLM基线的超越。

原文摘要 · Abstract (English)

Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evaluation of retrosynthesis performance remains limited. Existing benchmarks and metrics typically rely on published synthetic procedures and Top-K accuracy based on single ground-truth, which does not capture the open-ended nature of real-world synthesis planning. We propose a new benchmarking framework for single-step retrosynthesis that evaluates both general-purpose and chemistry-specialized LLMs using ChemCensor, a novel metric for chemical plausibility. By emphasizing plausibility over exact match, this approach better aligns with human synthesis planning practices. We also introduce CREED, a novel dataset comprising millions of ChemCensor-validated reaction records for LLM training, and use it to train a model that improves over the LLM baselines under this benchmark.

逆合成大模型评估药物发现化学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。