arXiv:2603.06722cs.LGcs.AI2026-03

用对比学习统一蛋白序列与结构表示,提升功能预测与结构检索

ProtAlign: Contrastive learning paradigm for Sequence and structure alignment

  • 构建序列与结构的对比对齐框架,学习跨模态共享嵌入空间
  • 在大规模序列-结构对上训练,显著提升功能注释与稳定性预测性能
  • 适合蛋白质工程、结构生物学研究者,支持跨模态检索与可解释分析

蛋白质语言模型通常关注序列与其文本描述之间的对齐,但忽略了结构信息。传统方法将序列与结构分开处理,限制了序列与结构嵌入之间对齐能力的发挥。本文提出一种序列-结构对比对齐框架,通过在大规模序列与实验解析或预测结构对上训练,使匹配的序列-结构对在嵌入空间中尽可能一致,不匹配对则被拉远。该对齐实现了跨模态检索(如根据序列找结构邻近物),提升了功能注释与稳定性估计等下游任务表现,并建立了序列变异与结构组织间的可解释联系。结果表明,对比学习可作为连接蛋白序列与结构的强大桥梁,提供统一表示以理解与设计蛋白质。

原文摘要 · Abstract (English)

Protein language models often take into consideration the alignment between a protein sequence and its textual description. However, they do not take structural information into consideration. Traditional methods treat sequence and structure separately, limiting the ability to exploit the alignment between the structure and protein sequence embeddings. In this paper, we introduce a sequence structure contrastive alignment framework, which learns a shared embedding space where proteins are represented consistently across modalities. By training on large-scale pairs of sequences and experimentally resolved or predicted structures, the model maximizes agreement between matched sequence structure pairs while pushing apart unrelated pairs. This alignment enables cross-modal retrieval (e.g., finding structural neighbors given a sequence), improves downstream prediction tasks such as function annotation and stability estimation, and provides interpretable links between sequence variation and structural organization. Our results demonstrate that contrastive learning can serve as a powerful bridge between protein sequences and structures, offering a unified representation for understanding and engineering proteins.

对比学习蛋白结构多模态表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。