arXiv:2510.23273cs.LGcs.AI2025-10

融合蛋白序列与结构信息,提升功能预测准确率。

A Novel Framework for Multi-Modal Protein Representation Learning

  • 用最优传输对齐不同模态嵌入空间,解决跨模态差异。
  • 条件图生成机制重构关系图,提升信息融合效果。
  • 理论+实证验证,适合生物信息学与多模态学习研究者。

精准的蛋白功能预测需整合序列、结构等内在信号与蛋白质互作、GO注释等外在噪声上下文。然而,两大挑战制约融合效果:(i) 预训练编码器生成的嵌入存在跨模态分布不一致;(ii) 外在数据关系图含噪声,损害GNN信息聚合。本文提出统一框架DAMPE,通过两项核心机制解决:首先,采用最优传输(OT)实现不同模态嵌入空间的对齐,缓解跨模态异质性;其次,设计条件图生成(CGG)方法,由条件编码器融合对齐后的内在嵌入,生成具信息量的图重建线索。理论分析表明,CGG目标促使编码器将图感知知识融入蛋白表示。实验上,DAMPE在标准GO基准上优于或匹配DPFunc,AUPR提升0.002–0.013个百分点,Fmax提升0.004–0.007个百分点。消融实验显示,OT对齐贡献AUPR提升0.043–0.064个百分点,CGG融合带来Fmax提升0.005–0.111个百分点。整体而言,DAMPE提供了一种可扩展且理论严谨的多模态蛋白表示学习方法,显著提升蛋白功能预测性能。

原文摘要 · Abstract (English)

Accurate protein function prediction requires integrating heterogeneous intrinsic signals (e.g., sequence and structure) with noisy extrinsic contexts (e.g., protein-protein interactions and GO term annotations). However, two key challenges hinder effective fusion: (i) cross-modal distributional mismatch among embeddings produced by pre-trained intrinsic encoders, and (ii) noisy relational graphs of extrinsic data that degrade GNN-based information aggregation. We propose Diffused and Aligned Multi-modal Protein Embedding (DAMPE), a unified framework that addresses these through two core mechanisms. First, we propose Optimal Transport (OT)-based representation alignment that establishes correspondence between intrinsic embedding spaces of different modalities, effectively mitigating cross-modal heterogeneity. Second, we develop a Conditional Graph Generation (CGG)-based information fusion method, where a condition encoder fuses the aligned intrinsic embeddings to provide informative cues for graph reconstruction. Meanwhile, our theoretical analysis implies that the CGG objective drives this condition encoder to absorb graph-aware knowledge into its produced protein representations. Empirically, DAMPE outperforms or matches state-of-the-art methods such as DPFunc on standard GO benchmarks, achieving AUPR gains of 0.002-0.013 pp and Fmax gains 0.004-0.007 pp. Ablation studies further show that OT-based alignment contributes 0.043-0.064 pp AUPR, while CGG-based fusion adds 0.005-0.111 pp Fmax. Overall, DAMPE offers a scalable and theoretically grounded approach for robust multi-modal protein representation learning, substantially enhancing protein function prediction.

蛋白功能预测多模态学习最优传输图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。