用图文对齐思路实现蛋白结构与功能的高效检索
Aligning Proteins and Language: A Foundation Model for Protein Retrieval
- 构建3D蛋白结构与功能描述的对比学习框架
- 在20万对数据上训练,跨数据库零样本检索表现优异
- 适合生物信息学与结构生物学研究者使用
本文旨在从大规模蛋白数据集中检索结构与功能相似的蛋白,辅助冷冻电镜等方法解析的蛋白结构的功能理解。受视觉语言模型进展启发,我们提出一种类似CLIP的框架,通过对比学习对齐3D蛋白结构与功能注释。为训练模型,我们构建了一个约20万对的蛋白-描述语句数据集,包含丰富的功能特征。我们在PDB和EMDB两个数据集上分别评估了域内与跨数据库检索性能,结果表明该方法在零样本场景下表现良好,展示了多模态基础模型在蛋白结构-功能理解中的潜力。
原文摘要 · Abstract (English)
This paper aims to retrieve proteins with similar structures and semantics from large-scale protein dataset, facilitating the functional interpretation of protein structures derived by structural determination methods like cryo-Electron Microscopy (cryo-EM). Motivated by the recent progress of vision-language models (VLMs), we propose a CLIP-style framework for aligning 3D protein structures with functional annotations using contrastive learning. For model training, we propose a large-scale dataset of approximately 200,000 protein-caption pairs with rich functional descriptors. We evaluate our model in both in-domain and more challenging cross-database retrieval on Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) dataset, respectively. In both cases, our approach demonstrates promising zero-shot retrieval performance, highlighting the potential of multimodal foundation models for structure-function understanding in protein biology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。