arXiv:2412.05269cs.LGcs.AI2024-12被引 9

通过融合多种模型偏差,提升逆合成预测准确率与化学家契合度。

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

  • 采用学习型集成框架融合两类具互补偏见的模型
  • 在小样本下仍保持高精度,且跨数据集零样本迁移成功
  • 工业化学家更偏好其预测结果,契合真实合成需求

化学合成仍是功能小分子发现与制造的关键瓶颈。基于AI的合成路径规划模型虽有进展,但仍难以处理罕见但关键的反应,且易产生幻觉或错误预测,影响多步搜索算法性能,并导致与化学家预期不一致。本文提出RetroChimera:一种前沿逆合成模型,基于两个具有互补归纳偏见的新组件,通过新型学习型集成框架融合多源预测。在多个数量级的数据规模和划分策略下实验表明,RetroChimera显著优于现有主流模型,展现出训练外数据的鲁棒性,首次实现每类反应仅需极少量样本即可学习。工业有机化学家对RetroChimera生成的反应路径评价更高,体现高度契合。此外,其在某大型制药公司内部数据集上实现零样本迁移,证明强泛化能力。该集成框架为构建更精准模型开辟新维度。

原文摘要 · Abstract (English)

Chemical synthesis remains a critical bottleneck in the discovery and manufacture of functional small molecules. AI-based synthesis planning models could be a potential remedy to find effective syntheses, and have made progress in recent years. However, they still struggle with less frequent, yet critical reactions for synthetic strategy, as well as hallucinated, incorrect predictions. This hampers multi-step search algorithms that rely on models, and leads to misalignment with chemists' expectations. Here we propose RetroChimera: a frontier retrosynthesis model, built upon two newly developed components with complementary inductive biases, which we fuse together using a new framework for integrating predictions from multiple sources via a learning-based ensembling strategy. Through experiments across several orders of magnitude in data scale and splitting strategy, we show RetroChimera outperforms all major models by a large margin, demonstrating robustness outside the training data, as well as for the first time the ability to learn from even a very small number of examples per reaction class. Moreover, industrial organic chemists prefer predictions from RetroChimera over the reactions it was trained on in terms of quality, revealing high levels of alignment. Finally, we demonstrate zero-shot transfer to an internal dataset from a major pharmaceutical company, showing robust generalization under distribution shift. With the new dimension that our ensembling framework unlocks, we anticipate further acceleration in the development of even more accurate models.

逆合成模型集成小样本学习化学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。