用最优传输生成伪标签,提升分子-蛋白互作预测准确率
KGOT: Unified Knowledge Graph and Optimal Transport Pseudo-Labeling for Molecule-Protein Interaction Prediction
- 基于最优传输构建跨模态伪标签生成机制
- 在多个数据集上显著提升预测精度与零样本能力
- 适合药物发现和计算生物学领域研究人员参考
分子-蛋白互作(MPI)预测是计算生物学中的基础任务,对药物研发和分子功能注释至关重要。现有模型面临两大挑战:一是标注数据稀少,真实互作仅占生物相关交互的一小部分;二是多数方法仅依赖分子和蛋白特征,忽略基因、代谢通路和功能注释等生物上下文信息。为此,本框架整合分子、蛋白、基因及通路级互作等多源生物数据,提出一种基于最优传输的伪标签生成方法,利用已知互作分布指导未标注对的标签分配。通过将伪标签视为连接异构生物模态的桥梁,有效融合多样化数据以增强MPI预测性能。在多个MPI数据集(包括虚拟筛选和蛋白检索任务)上的实验表明,该方法在预测准确率和零样本泛化能力方面均优于当前最优方法。本研究不仅推动了MPI预测发展,更提供了一种利用多元生物数据解决传统单/双模态学习限制的新范式,为计算生物学与药物发现带来新思路。
原文摘要 · Abstract (English)
Predicting molecule-protein interactions (MPIs) is a fundamental task in computational biology, with crucial applications in drug discovery and molecular function annotation. However, existing MPI models face two major challenges. First, the scarcity of labeled molecule-protein pairs significantly limits model performance, as available datasets capture only a small fraction of biological relevant interactions. Second, most methods rely solely on molecular and protein features, ignoring broader biological context such as genes, metabolic pathways, and functional annotations that could provide essential complementary information. To address these limitations, our framework first aggregates diverse biological datasets, including molecular, protein, genes and pathway-level interactions, and then develop an optimal transport-based approach to generate high-quality pseudo-labels for unlabeled molecule-protein pairs, leveraging the underlying distribution of known interactions to guide label assignment. By treating pseudo-labeling as a mechanism for bridging disparate biological modalities, our approach enables the effective use of heterogeneous data to enhance MPI prediction. We evaluate our framework on multiple MPI datasets including virtual screening tasks and protein retrieval tasks, demonstrating substantial improvements over state-of-the-art methods in prediction accuracies and zero shot ability across unseen interactions. Beyond MPI prediction, our approach provides a new paradigm for leveraging diverse biological data sources to tackle problems traditionally constrained by single- or bi-modal learning, paving the way for future advances in computational biology and drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。