arXiv:2412.08649q-bio.BMcs.LG2024-12被引 2

融合序列、文本与互作网络,提升小样本下的蛋白功能预测准确率

Multi-modal Representation Learning Enables Accurate Protein Function Prediction in Low-Data Setting

  • 整合蛋白质序列、生物文本和互作网络三模态信息构建综合表征
  • 在基因本体三大类别的基准测试中均优于现有方法
  • 适用于数据稀缺的生物研究场景,可发现潜在治疗靶点

本研究提出HOPER(HOlistic ProtEin Representation)这一新型多模态学习框架,旨在提升低数据环境下蛋白质功能预测(PFP)的准确性。蛋白质功能预测面临标注数据匮乏的挑战:传统机器学习模型在此情形下表现不佳,而深度学习模型虽在大数据下优异,但在数据稀缺时也难以奏效。HOPER通过融合蛋白质序列、生物医学文本及蛋白质-蛋白质互作(PPI)网络三种模态,利用自编码器生成全局嵌入表示,并结合迁移学习完成功能预测任务。在基因本体(GO)三大类别——分子功能、生物过程和细胞组分——的基准数据集上,HOPER均显著优于现有方法。此外,我们成功应用其识别出肺腺癌中的新型免疫逃逸蛋白,为潜在治疗靶点提供新线索。结果表明,多模态表示学习能有效缓解生物研究中的数据限制,推动更精准、可扩展的蛋白质功能预测。HOPER源代码与数据集已公开于https://github.com/kansil/HOPER。

原文摘要 · Abstract (English)

In this study, we propose HOPER (HOlistic ProtEin Representation), a novel multimodal learning framework designed to enhance protein function prediction (PFP) in low-data settings. The challenge of predicting protein functions is compounded by the limited availability of labeled data. Traditional machine learning models already struggle in such cases, and while deep learning models excel with abundant data, they also face difficulties when data is scarce. HOPER addresses this issue by integrating three distinct modalities - protein sequences, biomedical text, and protein-protein interaction (PPI) networks - to create a comprehensive protein representation. The model utilizes autoencoders to generate holistic embeddings, which are then employed for PFP tasks using transfer learning. HOPER outperforms existing methods on a benchmark dataset across all Gene Ontology categories, i.e., molecular function, biological process, and cellular component. Additionally, we demonstrate its practical utility by identifying new immune-escape proteins in lung adenocarcinoma, offering insights into potential therapeutic targets. Our results highlight the effectiveness of multimodal representation learning for overcoming data limitations in biological research, potentially enabling more accurate and scalable protein function prediction. HOPER source code and datasets are available at https://github.com/kansil/HOPER

蛋白功能预测多模态学习低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。