arXiv:2501.18439cs.LGq-bio.BM2025-01被引 1

用双尺度xLSTM提升分子表征,增强长程依赖建模能力。

MolGraph-xLSTM: A graph-based dual-level xLSTM framework with multi-head mixture-of-experts for enhanced molecular representation and interpretability

  • 分原子与基团两级处理分子图,结合xLSTM与跳跃知识捕获多层特征。
  • 在10个数据集上平均提升3.18% AUROC(分类)和3.83% RMSE(回归)。
  • 引入多头专家混合机制,兼顾性能与可解释性,适合药物发现场景。

分子性质预测对药物发现至关重要,计算方法可显著加速该过程。分子图已成为表示学习的焦点,图神经网络(GNNs)被广泛应用。然而,GNNs常难以捕捉长程依赖关系。为此,我们提出MolGraph-xLSTM,一种基于图的双尺度xLSTM模型,以增强特征提取并有效建模分子长程相互作用。该方法在原子级和基团级两个尺度处理分子图:原子级采用基于GNN的xLSTM框架结合跳跃知识,提取局部特征并聚合多层信息,有效捕捉局部与全局模式;基团级图提供更广泛的分子结构补充信息。两个尺度的嵌入经多头专家混合(MHMoE)机制优化,进一步提升表达能力与性能。我们在10个分子性质预测数据集上验证了MolGraph-xLSTM,涵盖分类与回归任务。模型在所有数据集上表现一致,分类任务中最高提升7.03%(BBBP数据集),回归任务中最高提升7.54%(ESOL数据集)。平均而言,相比基线方法,分类任务AUROC提升3.18%,回归任务RMSE降低3.83%。结果证实了模型的有效性,为药物发现中的分子表征学习提供了新方案。

原文摘要 · Abstract (English)

Predicting molecular properties is essential for drug discovery, and computational methods can greatly enhance this process. Molecular graphs have become a focus for representation learning, with Graph Neural Networks (GNNs) widely used. However, GNNs often struggle with capturing long-range dependencies. To address this, we propose MolGraph-xLSTM, a novel graph-based xLSTM model that enhances feature extraction and effectively models molecule long-range interactions. Our approach processes molecular graphs at two scales: atom-level and motif-level. For atom-level graphs, a GNN-based xLSTM framework with jumping knowledge extracts local features and aggregates multilayer information to capture both local and global patterns effectively. Motif-level graphs provide complementary structural information for a broader molecular view. Embeddings from both scales are refined via a multi-head mixture of experts (MHMoE), further enhancing expressiveness and performance. We validate MolGraph-xLSTM on 10 molecular property prediction datasets, covering both classification and regression tasks. Our model demonstrates consistent performance across all datasets, with improvements of up to 7.03% on the BBBP dataset for classification and 7.54% on the ESOL dataset for regression compared to baselines. On average, MolGraph-xLSTM achieves an AUROC improvement of 3.18\% for classification tasks and an RMSE reduction of 3.83\% across regression datasets compared to the baseline methods. These results confirm the effectiveness of our model, offering a promising solution for molecular representation learning for drug discovery.

分子表征xLSTM图神经网络药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。