融合蛋白-配体结构与基因序列,提升胃肠道疾病药物靶点预测准确率。
GastroDL-Fusion: A Dual-Modal Deep Learning Framework Integrating Protein-Ligand Complexes and Gene Sequences for Gastrointestinal Disease Drug Discovery
- 双模态架构:用图网络处理蛋白-配体结构,用预训练模型编码基因序列
- 在胃肠道疾病数据集上MAE达1.12,显著优于传统方法
- 适合药物研发人员用于靶点筛选和精准治疗设计
准确预测蛋白-配体结合亲和力对加速胃肠道疾病(如胃溃疡、克罗恩病、溃疡性结肠炎)新药与疫苗研发至关重要。传统计算模型仅依赖结构信息,难以捕捉影响疾病机制和治疗反应的遗传因素。为此,我们提出GastroDL-Fusion,一种融合蛋白-配体复合物与疾病相关基因序列信息的双模态深度学习框架。蛋白质-配体复合物以分子图表示,采用图同构网络(GIN)建模;基因序列通过预训练Transformer(ProtBERT/ESM)转化为生物学意义嵌入。两者通过多层感知机融合,实现跨模态交互学习。在胃肠道疾病相关靶标基准数据集上评估显示,该模型显著优于传统方法,达到均方误差(MAE)1.12,均方根误差(RMSE)1.75,超越了CNN、BiLSTM、GIN及仅使用Transformer的基线模型。结果表明,整合结构与遗传特征可更准确预测结合亲和力,为胃肠道疾病靶向疗法与疫苗设计提供可靠计算工具。
原文摘要 · Abstract (English)
Accurate prediction of protein-ligand binding affinity plays a pivotal role in accelerating the discovery of novel drugs and vaccines, particularly for gastrointestinal (GI) diseases such as gastric ulcers, Crohn's disease, and ulcerative colitis. Traditional computational models often rely on structural information alone and thus fail to capture the genetic determinants that influence disease mechanisms and therapeutic responses. To address this gap, we propose GastroDL-Fusion, a dual-modal deep learning framework that integrates protein-ligand complex data with disease-associated gene sequence information for drug and vaccine development. In our approach, protein-ligand complexes are represented as molecular graphs and modeled using a Graph Isomorphism Network (GIN), while gene sequences are encoded into biologically meaningful embeddings via a pre-trained Transformer (ProtBERT/ESM). These complementary modalities are fused through a multi-layer perceptron to enable robust cross-modal interaction learning. We evaluate the model on benchmark datasets of GI disease-related targets, demonstrating that GastroDL-Fusion significantly improves predictive performance over conventional methods. Specifically, the model achieves a mean absolute error (MAE) of 1.12 and a root mean square error (RMSE) of 1.75, outperforming CNN, BiLSTM, GIN, and Transformer-only baselines. These results confirm that incorporating both structural and genetic features yields more accurate predictions of binding affinities, providing a reliable computational tool for accelerating the design of targeted therapies and vaccines in the context of gastrointestinal diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。