用语言模型增强图网络,预测疾病与基因关联
GLaDiGAtor: Language-Model-Augmented Multi-Relation Graph Learning for Predicting Disease-Gene Associations
- 融合蛋白序列和疾病文本的上下文特征构建异质生物图
- 在14个基准方法中表现最优,新预测结果经文献验证
- 适合药物发现和基因功能研究者使用
理解疾病-基因关联对揭示疾病机制、推动诊断与治疗至关重要。传统人工梳理文献的方法耗时且难以扩展,促使人们采用机器学习处理大规模生物数据。特别是图神经网络(GNN)在建模复杂生物关系方面展现出潜力。为克服现有模型局限,我们提出GLaDiGAtor(基于图学习的疾病-基因关联预测),一种具有编码器-解码器架构的新GNN框架。该模型整合来自权威数据库的基因-基因、疾病-疾病及基因-疾病相互作用,构建异质生物图,并利用知名语言模型(ProtT5用于蛋白序列,BioBERT用于疾病文本)为每个节点注入上下文特征。在评估中,模型在预测准确性和泛化能力上均优于14种现有方法。文献支持的案例研究证实了高置信度新预测结果的生物学相关性,凸显其发现候选致病基因的潜力。这些结果证明图卷积网络在生物信息学中的强大能力,或可助力药物发现,揭示新的基因-疾病关联。源代码与处理后的数据集已公开于https://github.com/HUBioDataLab/GLaDiGAtor。
原文摘要 · Abstract (English)
Understanding disease-gene associations is essential for unravelling disease mechanisms and advancing diagnostics and therapeutics. Traditional approaches based on manual curation and literature review are labour-intensive and not scalable, prompting the use of machine learning on large biomedical data. In particular, graph neural networks (GNNs) have shown promise for modelling complex biological relationships. To address limitations in existing models, we propose GLaDiGAtor (Graph Learning-bAsed DIsease-Gene AssociaTiOn pRediction), a novel GNN framework with an encoder-decoder architecture for disease-gene association prediction. GLaDiGAtor constructs a heterogeneous biological graph integrating gene-gene, disease-disease, and gene-disease interactions from curated databases, and enriches each node with contextual features from well-known language models (ProtT5 for protein sequences and BioBERT for disease text). In evaluations, our model achieves superior predictive accuracy and generalisation, outperforming 14 existing methods. Literature-supported case studies confirm the biological relevance of high-confidence novel predictions, highlighting GLaDiGAtor's potential to discover candidate disease genes. These results underscore the power of graph convolutional networks in biomedical informatics and may ultimately facilitate drug discovery by revealing new gene-disease links. The source code and processed datasets are publicly available at https://github.com/HUBioDataLab/GLaDiGAtor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。