arXiv:2510.22008cs.LGq-bio.MN2025-10

构建了人类蛋白质多模态嵌入数据库,助力药物靶点发现。

A Multimodal Human Protein Embeddings Database: DeepDrug Protein Embeddings Bank (DPEB)

  • 整合四种蛋白嵌入:结构、序列、进化模式和词袋统计。
  • 图神经网络在蛋白互作预测上达87.37% AUROC,分类准确率超77%。
  • 适合系统生物学与药物研发人员使用,支持多种建模方法。

由于缺乏集成的多模态蛋白表示,计算预测蛋白-蛋白互作(PPI)面临挑战。DPEB 是一个精心整理的人类蛋白嵌入库,包含 22,043 个蛋白质,融合了四种嵌入类型:基于 AlphaFold2 的结构嵌入、基于 Transformer 的序列嵌入(BioEmbeddings)、基于 ESM-2 的上下文氨基酸模式嵌入,以及基于 n-gram 统计的序列嵌入(ProtVec)。尽管 AlphaFold2 蛋白结构可通过公开数据库获取,其内部神经网络嵌入仍不可得。DPEB 填补了这一空白,提供了 AlphaFold2 衍生的嵌入以支持计算建模。基准评估显示,结合 BioEmbedding 的 GraphSAGE 在 PPI 预测中表现最佳,达到 87.37% AUROC 与 79.16% 准确率;在酶分类任务中准确率为 77.42%,蛋白家族分类为 86.04%。DPEB 支持多种图神经网络方法,可用于系统生物学、药物靶点识别、通路分析及疾病机制研究。

原文摘要 · Abstract (English)

Computationally predicting protein-protein interactions (PPIs) is challenging due to the lack of integrated, multimodal protein representations. DPEB is a curated collection of 22,043 human proteins that integrates four embedding types: structural (AlphaFold2), transformer-based sequence (BioEmbeddings), contextual amino acid patterns (ESM-2: Evolutionary Scale Modeling), and sequence-based n-gram statistics (ProtVec]). AlphaFold2 protein structures are available through public databases (e.g., AlphaFold2 Protein Structure Database), but the internal neural network embeddings are not. DPEB addresses this gap by providing AlphaFold2-derived embeddings for computational modeling. Our benchmark evaluations show GraphSAGE with BioEmbedding achieved the highest PPI prediction performance (87.37% AUROC, 79.16% accuracy). The framework also achieved 77.42% accuracy for enzyme classification and 86.04% accuracy for protein family classification. DPEB supports multiple graph neural network methods for PPI prediction, enabling applications in systems biology, drug target identification, pathway analysis, and disease mechanism studies.

蛋白质嵌入多模态药物靶点图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。