arXiv:2502.04289physics.chem-phcs.LG2025-02被引 7

用排序模型预测无机材料合成路径,突破传统试错局限。

Retro-Rank-In: A Ranking-Based Approach for Inorganic Materials Synthesis Planning

  • 将目标与前驱物嵌入共享空间,通过成对排序选择最优合成路径。
  • 在未见数据上准确预测Cr2AlB2的前驱体组合CrB+Al,超越已有方法。
  • 适合需要高效设计新材料合成路线的研究者,尤其擅长新反应泛化。

逆合成旨在从简单易得的前驱物中规划目标化合物的合成路径,对新型无机材料合成至关重要,但传统方法仍依赖试错实验。现有机器学习方法因将逆合成视为多标签分类任务,难以泛化至全新反应。为此,我们提出Retro-Rank-In框架,将目标与前驱物嵌入共享潜在空间,并在无机化合物的二分图上学习成对排序器。我们在针对数据重复和重叠问题设计的挑战性数据集上评估其泛化能力。例如,对于Cr2AlB2,模型成功预测出从未在训练中出现过的验证前驱体组合CrB + Al,而此前方法无法实现此能力。大量实验表明,Retro-Rank-In在分布外泛化与候选集排序方面达到新基准,显著加速无机材料合成进程。

原文摘要 · Abstract (English)

Retrosynthesis strategically plans the synthesis of a chemical target compound from simpler, readily available precursor compounds. This process is critical for synthesizing novel inorganic materials, yet traditional methods in inorganic chemistry continue to rely on trial-and-error experimentation. Emerging machine-learning approaches struggle to generalize to entirely new reactions due to their reliance on known precursors, as they frame retrosynthesis as a multi-label classification task. To address these limitations, we propose Retro-Rank-In, a novel framework that reformulates the retrosynthesis problem by embedding target and precursor materials into a shared latent space and learning a pairwise ranker on a bipartite graph of inorganic compounds. We evaluate Retro-Rank-In's generalizability on challenging retrosynthesis dataset splits designed to mitigate data duplicates and overlaps. For instance, for Cr2AlB2, it correctly predicts the verified precursor pair CrB + Al despite never seeing them in training, a capability absent in prior work. Extensive experiments show that Retro-Rank-In sets a new state-of-the-art, particularly in out-of-distribution generalization and candidate set ranking, offering a powerful tool for accelerating inorganic material synthesis.

逆合成材料发现排序模型机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。