arXiv:2602.13136cs.LG2026-02被引 1

通过原子顺序设计提升逆合成生成效率与精度

Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching

  • 以反应中心原子前置构建位置先验,显式编码化学知识
  • 在USPTO-50k上达61.2%顶1准确率,生成仅需20-50步
  • 小模型加结构先验胜过大模型无序训练,适合高效研发

无模板逆合成方法将任务视为黑箱序列生成,学习效率受限;半模板方法依赖僵化反应库,泛化能力差。本文提出关键洞见:神经表示中的原子顺序至关重要。基于此,我们设计结构感知的无模板框架,将化学反应的两阶段特性作为位置归纳偏置。通过将反应中心原子置于序列开头,方法将隐含化学知识转化为模型可捕捉的显式位置模式。所提RetroDiT骨干为带旋转位置嵌入的图变压器,利用此顺序优先关注化学关键区域。结合离散流匹配,实现训练与采样解耦,生成步数降至20–50步(此前扩散方法需500步)。在USPTO-50k(顶1准确率61.2%)和USPTO-Full(51.3%)上达当前最优,使用已知反应中心时分别达71.1%与63.4%,超越训练于100亿反应的基础模型,且数据量仅为其万分之一。消融实验表明结构先验优于暴力扩展:28万参数模型有序时性能媲美6500万参数无序模型。

原文摘要 · Abstract (English)

Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization. We address this gap with a key insight: atom ordering in neural representations matters. Building on this insight, we propose a structure-aware template-free framework that encodes the two-stage nature of chemical reactions as a positional inductive bias. By placing reaction center atoms at the sequence head, our method transforms implicit chemical knowledge into explicit positional patterns that the model can readily capture. The proposed RetroDiT backbone, a graph transformer with rotary position embeddings, exploits this ordering to prioritize chemically critical regions. Combined with discrete flow matching, our approach decouples training from sampling and enables generation in 20--50 steps versus 500 for prior diffusion methods. Our method achieves state-of-the-art performance on both USPTO-50k (61.2% top-1) and the large-scale USPTO-Full (51.3% top-1) with predicted reaction centers. With oracle centers, performance reaches 71.1% and 63.4% respectively, surpassing foundation models trained on 10 billion reactions while using orders of magnitude less data. Ablation studies further reveal that structural priors outperform brute-force scaling: a 280K-parameter model with proper ordering matches a 65M-parameter model without it.

逆合成生成模型图神经网络化学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。