融合化学扰动与结构信息,提升药物靶点亲和力预测精度
GramSeq-DTA: A grammar-based drug-target affinity prediction approach fusing gene expression information
- 用语法变分自编码器提取药物结构特征,结合基因表达扰动数据
- 在三个公开数据集上超越现有最优模型,提升预测准确率
- 适合关注多模态生物信息融合的药物研发研究人员
药物-靶点亲和力(DTA)预测是药物发现的关键环节。现有基于一维字符串的表示方法虽有效,但忽略了原子和键的相对位置信息。图结构表示虽部分解决此问题,但仅考虑结构特征仍不足以精确预测。本文提出GramSeq-DTA,融合药物的化学扰动信息与结构特征。采用语法变分自编码器(GVAE)提取药物特征,使用卷积神经网络(CNN)和循环神经网络(RNN)分别处理蛋白质序列特征。化学扰动数据来自L1000项目,包含药物引起基因的上调与下调信息,经处理后作为药物的功能特征。通过整合药物、基因与靶点特征,模型在BindingDB、Davis和KIBA等常用数据集上优于当前最先进方法,验证了多模态融合的有效性,为DTA预测提供了新范式。
原文摘要 · Abstract (English)
Drug-target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug-target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance the prediction capability of the models. To fill this gap, we propose GramSeq-DTA, which integrates chemical perturbation information with the structural information of drugs and targets. We applied a Grammar Variational Autoencoder (GVAE) for drug feature extraction and utilized two different approaches for protein feature extraction: Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN). The chemical perturbation data is obtained from the L1000 project, which provides information on the upregulation and downregulation of genes caused by selected drugs. This chemical perturbation information is processed, and a compact dataset is prepared, serving as the functional feature set of the drugs. By integrating the drug, gene, and target features in the model, our approach outperforms the current state-of-the-art DTA prediction models when validated on widely used DTA datasets (BindingDB, Davis, and KIBA). This work provides a novel and practical approach to DTA prediction by merging the structural and functional aspects of biological entities, and it encourages further research in multi-modal DTA prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。