用多答案评估和训练,让大模型更懂化学反应的多样性。
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

- 采用Top-K提示法,同时生成多个合理反应路径
- 在4560万条验证反应数据上训练出新模型,性能领先
- 适合需要多样化合成方案的研发人员
单步逆合成是计算机辅助合成规划的核心,但其固有的多对一特性难以通过单一答案的评估方式准确捕捉。为此,我们提出Top-K提示法作为稳健的训练与推理范式,以更好反映多样且合理的反应预测。我们构建了包含约4560万条已验证反应的超大规模数据集CREED-CCV-2+USPTO-XL,用于训练C3LM(Chemistry Constraint-Consistent Language Model)。通过结合微调与基于ChemCensor的可塑性奖励及新颖性奖励,该模型在OOD URSA-expert-2026基准上达到当前最优性能。进一步分析显示,大语言模型与传统模型探索的反应空间具有互补性,这推动了集成式逆合成系统的发展。总体而言,我们的研究确立了Top-K与可塑性感知训练为未来基于大模型的合成规划中实用的新方向。
原文摘要 · Abstract (English)
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。