提出新分子表示法,让逆合成预测更精准高效
Copy-Augmented Representation for Structure Invariant Template-Free Retrosynthesis
- 用五种特殊标记分解SMILES,缩小反应物与产物差异
- 99.9%生成分子有效,USPTO-50K上准确率达67.2%
- 适合药物发现中的结构敏感分子生成任务
逆合成预测是药物发现与化学合成的核心,需识别能生成目标分子的反应物。现有无模板方法难以捕捉化学反应中大量分子骨架保持不变的结构不变性,导致搜索空间过大、预测准确率下降。本文提出C-SMILES,将传统SMILES分解为元素-标记对,引入五个特殊标记,显著降低反应物与产物间的编辑距离。基于此表示,设计复制增强机制,动态判断是否生成新原子或保留产物中未变化的片段。结合SMILES对齐引导,提升注意力与真实原子映射的一致性,实现更符合化学逻辑的预测。在USPTO-50K和大规模USPTO-FULL数据集上的全面评估显示:在USPTO-50K上达到67.2%的top-1准确率,在USPTO-FULL上达50.8%,生成分子有效性高达99.9%。该工作建立了一种面向结构感知的分子生成新范式,可直接应用于计算药物发现。
原文摘要 · Abstract (English)
Retrosynthesis prediction is fundamental to drug discovery and chemical synthesis, requiring the identification of reactants that can produce a target molecule. Current template-free methods struggle to capture the structural invariance inherent in chemical reactions, where substantial molecular scaffolds remain unchanged, leading to unnecessarily large search spaces and reduced prediction accuracy. We introduce C-SMILES, a novel molecular representation that decomposes traditional SMILES into element-token pairs with five special tokens, effectively minimizing editing distance between reactants and products. Building upon this representation, we incorporate a copy-augmented mechanism that dynamically determines whether to generate new tokens or preserve unchanged molecular fragments from the product. Our approach integrates SMILES alignment guidance to enhance attention consistency with ground-truth atom mappings, enabling more chemically coherent predictions. Comprehensive evaluation on USPTO-50K and large-scale USPTO-FULL datasets demonstrates significant improvements: 67.2% top-1 accuracy on USPTO-50K and 50.8% on USPTO-FULL, with 99.9% validity in generated molecules. This work establishes a new paradigm for structure-aware molecular generation with direct applications in computational drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。