arXiv:2409.01731cs.LG2024-09被引 1

融合分子文本与图结构的多模态模型,提升毒性预测准确率

Stacked ensemble\-based mutagenicity prediction model using multiple modalities with graph attention network

  • 用SMILES和分子图双模态提取分子特征
  • 堆叠集成模型在Hansen数据集上达95.21%的AUC
  • 结合SHAP解释模型,适合药物研发与计算生物学家

致突变性因与基因突变相关,可能引发癌症等负面后果,早期识别药物中的致突变化合物对避免无效候选物、降低研发成本至关重要。现有计算方法多依赖单一模态,本研究提出一种基于堆叠集成的多模态致突变性预测模型,融合简化分子输入线性输入系统(SMILES)与分子图。其中,SMILES用于提取子结构、理化及几何信息,分子图通过图注意力网络(GAT)捕捉拓扑特征。模型采用机器学习分类器堆叠集成进行预测,并使用可解释人工智能技术SHAP分析各分类器重要性及关键特征。实验表明,该方法在两个标准数据集上优于现有最先进模型,尤其在Hansen基准数据集上达到95.21%的受试者工作特征曲线下面积(AUC),验证了其有效性。该研究对转化医学领域的临床医生与计算生物学家具有吸引力。

原文摘要 · Abstract (English)

Mutagenicity is a concern due to its association with genetic mutations which can result in a variety of negative consequences, including the development of cancer. Earlier identification of mutagenic compounds in the drug development process is therefore crucial for preventing the progression of unsafe candidates and reducing development costs. While computational techniques, especially machine learning models have become increasingly prevalent for this endpoint, they rely on a single modality. In this work, we introduce a novel stacked ensemble based mutagenicity prediction model which incorporate multiple modalities such as simplified molecular input line entry system (SMILES) and molecular graph. These modalities capture diverse information about molecules such as substructural, physicochemical, geometrical and topological. To derive substructural, geometrical and physicochemical information, we use SMILES, while topological information is extracted through a graph attention network (GAT) via molecular graph. Our model uses a stacked ensemble of machine learning classifiers to make predictions using these multiple features. We employ the explainable artificial intelligence (XAI) technique SHAP (Shapley Additive Explanations) to determine the significance of each classifier and the most relevant features in the prediction. We demonstrate that our method surpasses SOTA methods on two standard datasets across various metrics. Notably, we achieve an area under the curve of 95.21\% on the Hansen benchmark dataset, affirming the efficacy of our method in predicting mutagenicity. We believe that this research will captivate the interest of both clinicians and computational biologists engaged in translational research.

毒性预测多模态图神经网络SHAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。