arXiv:2606.12609cs.LGq-bio.QM2026-06中稿 · ICML

用病毒蛋白揭示语言模型如何编码生物序列的天然性

Viral Proteins Reveal Geometry of Protein Language Models

论文配图:Viral Proteins Reveal Geometry of Protein Language Models
图 1 · 摘自论文原文
  • 以病毒蛋白为探针,发现嵌入空间存在天然性主轴
  • 病毒蛋白在模型中仍保持可区分性,超越零样本困惑度
  • 适合研究模型对稀有生物序列的表征能力

蛋白质语言模型在高度不平衡的数据集上训练,引发其对罕见生物序列表示能力的疑问。本文以不同ESM模型家族中的病毒蛋白为例,发现嵌入空间中存在一条主导的天然性轴,该轴与掩码重建困惑度对齐,将序列从典型细胞蛋白、病毒蛋白排序至随机序列。随着模型规模增大,这一轴在不同病毒家族间收缩不均。尽管如此,模型嵌入仍保留病毒特异性信号:病毒蛋白在零样本困惑度和浅层序列特征之外仍可线性分离。结果表明,蛋白质语言模型的表示既受普遍天然性概念支配,也保留了特定生物类群的信息。

原文摘要 · Abstract (English)

Protein language models are trained on highly imbalanced datasets, raising the question of how they represent underrepresented biological sequences. Using viral proteins as a case study across ESM model families, we identify a dominant nativeness axis in embedding space, aligned with masked reconstruction perplexity, that orders sequences from well-modeled cellular proteins through viral proteins to shuffled and random sequences. Scaling contracts this axis unevenly across viral families. Despite this, protein language model embeddings retain viral-specific signal: viral proteins remain linearly separable beyond zero-shot perplexity and shallow sequence features. Together, these results suggest that pLM representations are structured by a general notion of nativeness while preserving information specific to distinct biological groups.

蛋白质语言模型嵌入空间病毒蛋白

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。