融合文本与生物序列信息,提升医学知识图谱链接预测能力
Multimodal Contrastive Representation Learning in Augmented Biomedical Knowledge Graphs
- 用语言模型与图对比学习融合多模态嵌入,增强实体内部关系表征
- 在PrimeKG++和DrugBank数据集上实现高精度链接预测,对未见节点也有效
- 适合从事药物发现、生物信息学的研究者参考
生物医学知识图谱(BKGs)整合多种数据以揭示复杂的生物医学关系。有效的链接预测可发现潜在的新药-疾病关联。本文提出一种新型多模态方法,将专用语言模型(LMs)的嵌入与图对比学习(GCL)结合,强化实体内关系表征,并利用知识图谱嵌入(KGE)模型捕捉实体间关系,实现高效链接预测。为克服现有BKGs的局限性,我们构建了包含生物序列和文本描述的增强型知识图谱PrimeKG++。通过统一语义与关系信息,该方法展现出强泛化能力,在未见节点上也能实现准确预测。在PrimeKG++和DrugBank药物-靶点互作数据集上的实验验证了方法的有效性与鲁棒性。源代码、预训练模型及数据已公开于https://github.com/HySonLab/BioMedKG。
原文摘要 · Abstract (English)
Biomedical Knowledge Graphs (BKGs) integrate diverse datasets to elucidate complex relationships within the biomedical field. Effective link prediction on these graphs can uncover valuable connections, such as potential novel drug-disease relations. We introduce a novel multimodal approach that unifies embeddings from specialized Language Models (LMs) with Graph Contrastive Learning (GCL) to enhance intra-entity relationships while employing a Knowledge Graph Embedding (KGE) model to capture inter-entity relationships for effective link prediction. To address limitations in existing BKGs, we present PrimeKG++, an enriched knowledge graph incorporating multimodal data, including biological sequences and textual descriptions for each entity type. By combining semantic and relational information in a unified representation, our approach demonstrates strong generalizability, enabling accurate link predictions even for unseen nodes. Experimental results on PrimeKG++ and the DrugBank drug-target interaction dataset demonstrate the effectiveness and robustness of our method across diverse biomedical datasets. Our source code, pre-trained models, and data are publicly available at https://github.com/HySonLab/BioMedKG
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。