提出Align-React框架,精准捕捉反应物到产物的原子对应关系。
A General-Purpose Framework for Chemical Reaction Representation with Atomic Correspondence and Flexible Condition Adaptation
- 通过原子对应关系建模分子转化过程,提升结构理解精度
- 在多个基准数据集上超越现有模型,性能显著提升
- 适配不同反应条件,适合药物研发等复杂合成任务
有机合成是化学工业的基础,尤其在制药领域至关重要。尽管人工智能为化学反应建模提供了强大工具,但当前方法主要局限于两类:依赖人工设计的领域特异性特征,或通过简单拼接/聚合反应组分的通用深度学习模型。前者难以随数据量扩展,后者因简单拼接无法直接捕捉反应物与产物间的结构变化,且难以适应含非分子反应条件的数据集而无需修改架构。本文提出Align-React,一种通用化学反应表征学习框架,通过整合反应物与产物间的原子对应关系,精确识别分子转化模式,增强对转化规律的理解。引入适配器结构嵌入反应条件,提升跨数据集和任务的适应性。同时提出反应中心感知注意力机制,使模型聚焦关键官能团,生成更强大、更具信息量的表征。在多个下游任务中评估,该模型在多数基准数据集上显著优于现有化学反应表征学习架构。
原文摘要 · Abstract (English)
Motivation: Organic synthesis is fundamental to the chemical industry, particularly in domains such as pharmaceutical development. While artificial intelligence offers powerful tools for modeling chemical reactions, current approaches are primarily limited to two paradigms: those that rely on hand-crafted, domain-specific features, and those that apply generic deep learning models through simplistic concatenation or aggregation of reaction components. The former often struggles to scale effectively with increasing data volumes, while the latter relies on simple input- or feature-level concatenation to combine different reaction components. Such simplistic aggregation prevents these models from directly capturing the structural transformations between reactants and products, and also makes it difficult to adapt or extend them to datasets that include non-molecular reaction conditions without modifying their model architectures. Results: This paper introduces Align-React, a novel chemical reaction representation learning framework designed for a wide range of organic reaction tasks. Our approach integrates atomic correspondence between reactants and products to discern precise molecular transformations, thereby improving the model's comprehension of molecular transformation patterns. We incorporate an adapter structure to embed reaction conditions into the representation, enhancing adaptability across varied datasets and tasks. Furthermore, a Reaction-Center-Aware attention mechanism is proposed to enable the model to focus on critical functional groups, yielding more powerful and informative representations. Evaluated across multiple downstream tasks, our model demonstrates superior performance, significantly outperforming existing chemical reaction representation learning architectures on most benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。