用新方法提升化学合成路径规划的准确率和多样性。
Margin-calibrated Classifier Guidance for Property-driven Synthesis Planning

- 提出SCR方法,通过对比学习和边缘校准优化分类器。
- 在USPTO-190上,引导下求解率从16.8%提升至95.3%。
- 适合需要精准控制反应类型或分子相似性的化学研究者。
合成规划旨在找到高效生成目标分子的反应序列。通常使用预训练的单步自回归逆合成模型反复调用以生成序列。分类器引导理论上可在不重训练生成器的情况下,帮助模型输出满足特定约束或符合化学家偏好的反应。我们发现,仅用交叉熵损失训练的辅助分类器无法有效覆盖无条件的词级分布。为此,我们提出一种新方法——序列补全排序(SCR),结合对比论证与基于边缘的损失,校准分类器以在解码过程中有效区分不同延续路径。我们形式化证明了边缘校准分类器可扩展引导束搜索下可达的属性满足序列集合。实验表明,在USPTO-190数据集上,给定化学家指定的引导目标,SCR将多步求解率从无引导生成器的16.8%大幅提升至反应类型引导下的78.4%和Tanimoto相似性引导下的95.3%,解锁了33个此前基线无法解决的目标(占17.4%)。该方法还有效弥合了无模板与有模板方法间的长期多样性差距。
原文摘要 · Abstract (English)
Synthesis planning seeks an efficient sequence of chemical reactions that produce a target molecule. Typically, a pretrained single-step (autoregressive) retrosynthesis model is repeatedly invoked to generate such a sequence. Classifier guidance can, in principle, help steer the output of single-step model toward reactions that satisfy specific constraints or accommodate chemist's preferences during inference without having to retrain the autoregressive generator. We expose the insufficiency of auxiliary classifiers trained with cross-entropy loss to override the unconditional token-level distributions learned from typical sparse single-disconnection reaction datasets. We overcome this issue with a novel method called Sequence Completion Ranking (SCR), which employs contrastive argumentation and a margin-based loss to calibrate the classifier so that it can meaningfully discriminate between continuations during decoding. We formally establish that margin-calibrated classifiers can expand the set of property-satisfying sequences reachable under guided beam search. Empirically, on USPTO-190, given chemist-specified guidance targets, SCR substantially improves multi-step solve rates from $16.8\%$ (unguided generator) to $78.4\%$ with reaction-type guidance and $95.3\%$ with Tanimoto guidance, unlocking valid routes for 33 targets ($17.4\%$) previously unsolvable with baselines. Our method also effectively closes the long-standing diversity gap between template-free and template-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。