arXiv:2608.18940cs.LGcs.AI2026-08

用多答案评估和训练,让大模型更懂化学反应的多样性。

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

论文配图:Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
图 1 · 摘自论文原文
  • 采用Top-K提示法,同时生成多个合理反应路径
  • 在4560万条验证反应数据上训练出新模型,性能领先
  • 适合需要多样化合成方案的研发人员

单步逆合成是计算机辅助合成规划的核心,但其固有的多对一特性难以通过单一答案的评估方式准确捕捉。为此,我们提出Top-K提示法作为稳健的训练与推理范式,以更好反映多样且合理的反应预测。我们构建了包含约4560万条已验证反应的超大规模数据集CREED-CCV-2+USPTO-XL,用于训练C3LM(Chemistry Constraint-Consistent Language Model)。通过结合微调与基于ChemCensor的可塑性奖励及新颖性奖励,该模型在OOD URSA-expert-2026基准上达到当前最优性能。进一步分析显示,大语言模型与传统模型探索的反应空间具有互补性,这推动了集成式逆合成系统的发展。总体而言,我们的研究确立了Top-K与可塑性感知训练为未来基于大模型的合成规划中实用的新方向。

原文摘要 · Abstract (English)

Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

逆合成大模型化学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。