MOL-Mamba融合结构与电子信息,提升分子表征能力
MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic Insights
- 分层结构推理+电子关联融合,双路径建模分子特性
- 在11个数据集上超越现有方法,平均性能提升显著
- 适合药物设计、分子性质预测等化学人工智能任务
分子表征学习在分子性质预测和药物设计等下游任务中至关重要。图神经网络(GNN)和图变换器(GT)在自监督预训练方面展现出潜力,但现有方法常忽略分子结构与电子信息之间的关系,以及分子内部的语义推理过程。这种对基本化学知识的缺失导致表征不完整,未能整合结构与电子数据。为此,我们提出MOL-Mamba框架,通过原子与片段级分层结构推理(MG模块)和结构-电子关联融合(MT模块),实现对分子结构与电子特性的联合建模。此外,我们设计了结构分布协同训练与电子语义融合训练策略,进一步提升表征能力。大量实验表明,MOL-Mamba在11个化学-生物分子数据集上均优于当前最优基线。
原文摘要 · Abstract (English)
Molecular representation learning plays a crucial role in various downstream tasks, such as molecular property prediction and drug design. To accurately represent molecules, Graph Neural Networks (GNNs) and Graph Transformers (GTs) have shown potential in the realm of self-supervised pretraining. However, existing approaches often overlook the relationship between molecular structure and electronic information, as well as the internal semantic reasoning within molecules. This omission of fundamental chemical knowledge in graph semantics leads to incomplete molecular representations, missing the integration of structural and electronic data. To address these issues, we introduce MOL-Mamba, a framework that enhances molecular representation by combining structural and electronic insights. MOL-Mamba consists of an Atom & Fragment Mamba-Graph (MG) for hierarchical structural reasoning and a Mamba-Transformer (MT) fuser for integrating molecular structure and electronic correlation learning. Additionally, we propose a Structural Distribution Collaborative Training and E-semantic Fusion Training framework to further enhance molecular representation learning. Extensive experiments demonstrate that MOL-Mamba outperforms state-of-the-art baselines across eleven chemical-biological molecular datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。