arXiv:2607.01627cs.LGcs.AI2026-07

用多模态知识图谱提升冷启动蛋白互作预测准确率

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

论文配图:MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction
图 1 · 摘自论文原文
  • 融合序列与四大生物知识图谱,捕捉蛋白关联上下文
  • 在新蛋白-旧蛋白和新蛋白-新蛋白场景下,各项指标均领先
  • 适合药物研发和功能基因组学中的新靶点发现

准确预测蛋白-蛋白相互作用(PPI)对功能基因组学、疾病机制研究和药物开发至关重要。当候选互作涉及训练中未出现的蛋白质时,仅依赖网络拓扑的模型常丢失有效信息。本文提出 extit{MKGR},一种用于冷启动PPI预测的多模态表征框架。该方法结合区域感知的蛋白序列编码与四种以蛋白为中心的生物医学知识图谱:蛋白-药物、蛋白-疾病、蛋白-miRNA及蛋白-lncRNA关联。序列分支从结构信息引导的序列区域提取上下文表征,图注意力编码器则从稀疏的生物医学关联中学习模态特异的蛋白嵌入。桥接重建目标通过恢复共享的蛋白-实体关联来正则化图学习,成对门控模块自适应整合每对候选蛋白的序列与图证据。在两个基准数据集上,于新-旧和新-新冷启动设置下, extit{MKGR}在准确率(ACC)、F1、AUC、AUPR和MCC等指标上持续优于序列、网络及知识图谱基线模型。

原文摘要 · Abstract (English)

Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficult setting arises when candidate interactions include proteins that have no observed PPI edges during training, where models relying on network topology alone often lose useful context. This paper presents \method, a multimodal representation framework for cold-start PPI prediction. \method\ combines region-aware protein sequence encoding with four protein-centered biomedical knowledge graphs, including protein-drug, protein-disease, protein-miRNA, and protein-lncRNA associations. The sequence branch extracts contextual representations from structurally informed sequence regions, while graph attention encoders learn modality-specific protein embeddings from sparse biomedical associations. A bridge reconstruction objective regularizes graph learning by recovering shared protein-entity associations, and a pair-level gating module adaptively integrates sequence and graph evidence for each candidate protein pair. Experiments on two benchmark datasets under novel-old and novel-novel cold-start settings show that \method\ consistently outperforms competitive sequence, network, and knowledge-graph baselines across ACC, F1, AUC, AUPR, and MCC.

蛋白互作冷启动多模态知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。