arXiv:2605.29158cs.LGcs.IR2026-05

用残基级匹配提升远缘蛋白同源搜索精度

PROTOCOL: Late Interaction Retrieval for Protein Homolog Search

论文配图:PROTOCOL: Late Interaction Retrieval for Protein Homolog Search
图 1 · 摘自论文原文
  • 保留蛋白残基嵌入,通过晚期交互比较相似性
  • 在SCOPe和Pfam数据集上超越现有方法
  • 适合需要精准远缘同源识别的研究者

蛋白质同源搜索对功能注释、结构预测和进化分析至关重要,但在序列相似性低的'黄昏区'仍具挑战性。传统比对方法在此区域灵敏度下降。蛋白质语言模型虽能提供上下文感知表示,但以往基于嵌入的检索流程常将表示池化为单一向量,可能丢失局部保守模体或功能残基。本文提出ProtoCol模型,将蛋白表示为残基嵌入集合,采用类似ColBERT的晚期交互机制,通过残基级对比提升同源检索能力。该模型独立编码蛋白,候选物表示可预计算,使用最大相似度(MaxSim)进行打分。在SCOPe超家族与Pfam族基准测试中,ProtoCol优于序列组成、比对方法、池化蛋白语言模型及单向量训练基线,验证了晚期交互作为远程同源搜索有效检索层的潜力。

原文摘要 · Abstract (English)

Protein homology search underlies function annotation, structure prediction, and evolutionary analysis, but remains challenging in the "twilight zone," where global sequence similarity is weak and classical alignment methods lose sensitivity. Protein language models provide context-aware representations that could improve alignment sensitivity in this regime. However, prior protein embedding-based retrieval pipelines often pool these representations into a single vector, potentially obscuring local motifs, domains, or conserved residues that reveal remote homology. We introduce ProtoCol, a model which represents proteins as sets of residue embeddings and uses ColBERT-style late interaction to test whether residue-level comparison improves homolog retrieval. ProtoCol encodes proteins independently, keeps candidate representations pre-computable, and scores candidates with MaxSim over residue embeddings. On SCOPe superfamily and Pfam clan benchmarks, ProtoCol outperforms sequence-composition, alignment-based, pooled PLM, and trained single-vector baselines, supporting late interaction as an effective retrieval layer for remote homology search.

蛋白搜索同源识别嵌入匹配延迟交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。