对比多种分子编码方法,提升药物属性预测准确率与可解释性。
A systematic investigation of molecular encoding methods for drug property predictions across neural network and Transformer encoder-based model
- 用MLP和Transformer模型测试指纹、子结构、字符串等编码方式。
- 多任务分类平均AUC超0.9,毒性、致突变性预测效果佳。
- 通过注意力权重直接识别关键化学基团,适合药物研发人员参考。
针对不同分子编码方法对分子属性预测的影响,本研究系统评估了两种主流模型架构:经典全连接网络(MLP)与基于Transformer编码器的模型(MLP+TL)。在七个知名分子数据集上,考察了拓扑指纹、子结构指纹及字符串表示等多种编码方法。在包括毒性、致突变性和副作用预测在内的多个生物相关分类任务中,模型平均AUC值均超过0.9。不同于依赖LIME或SHAP等外部解释方法,本研究利用模型内部注意力权重作为可解释信号,识别出决定血脑屏障通透性及沙门氏菌致突变性的化学可解释特征。以吗啡与海洛因为例,注意力权重一致反映出羟基相关亚结构在血脑屏障通透性中的关键作用。研究为高效分子编码选择提供了实证指导,并推动可解释分子信息学在药物发现中的应用。
原文摘要 · Abstract (English)
Fundamental investigations into how different molecular encoding methods affect molecular property prediction remain relatively limited. In this study, we extensively examined the optimal molecular encoding methods for molecular properties prediction using two prevalent structure designs: a classical neural network model (MLP) and a Transformer encoder-based model (MLP+TL). For molecular encoding methods, we investigated several types of fingerprints, including traditional topological fingerprints, substructure-based fingerprints, and string-based representations. These two models were trained on seven well-known molecular datasets to evaluate different input molecular encoding methods based on evaluation metrics. On several biologically relevant classification tasks, including toxicity, mutagenicity, and side-effect prediction, our models consistently achieved average AUC values above 0.9. Rather than relying on external post-hoc explanation methods such as the local interpretable model-agnostic explanation (LIME) or the Deep SHapley Additive exPlanations (SHAP), we leveraged the model's intrinsic attention weights as an internal interpretability signal for identifying potentially important feature. The MLP+TL model using MACCS and PubChem as input can capture chemically interpretable groups that determined the major blood-brain barrier (BBB) permeability and mutagenicity in Salmonella typhimurium. In particular, a comparison between Morphine and Heroin highlighted the role of hydroxyl-related substructures in BBB permeability prediction, which was consistently reflected in the attention weights. Overall, our findings provide practical guidance for selecting effective molecular encoding methods and contribute to the development of interpretable molecular informatics approaches for drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。