arXiv:2509.24895cs.LG2025-09中稿 · ICLR被引 2

用形状分析法揭示蛋白语言模型的表征演化规律

Towards Understanding the Shape of Representations in Protein Language Models

  • 引入SRV与图过滤技术构建蛋白表征度量空间
  • 发现ESM2各层表征的均值与维度呈非线性变化
  • 适合研究模型表征机制与结构预测的科研人员

尽管蛋白语言模型(PLMs)在从头蛋白设计中前景广阔,但其如何将序列转化为隐藏表征,以及这些表征中编码的信息仍不明确。现有可解释性工具多关注单个序列的变换,而对整个序列空间及其关系的转化尚不清楚。本文通过将蛋白结构与表征映射为平方根速度(SRV)表示及图过滤方法,构建了可用于比较蛋白或其表征的度量空间。基于SCOP数据集分析不同蛋白,发现ESM2模型不同层数的Karcher均值和有效维度随层数呈现非线性模式。同时,利用图过滤研究模型在不同上下文长度下对蛋白结构特征的编码能力,结果显示模型更擅长捕捉残基间的局部及邻近关系,超过一定上下文长度后性能下降。最结构忠实的编码通常出现在接近但早于最后一层的位置,提示在这些层上训练折叠模型可能提升折叠效果。

原文摘要 · Abstract (English)

While protein language models (PLMs) are one of the most promising avenues of research for future de novo protein design, the way in which they transform sequences to hidden representations, as well as the information encoded in such representations is yet to be fully understood. Several works have attempted to propose interpretability tools for PLMs, but they have focused on understanding how individual sequences are transformed by such models. Therefore, the way in which PLMs transform the whole space of sequences along with their relations is still unknown. In this work we attempt to understand this transformed space of sequences by identifying protein structure and representation with square-root velocity (SRV) representations and graph filtrations. Both approaches naturally lead to a metric space in which pairs of proteins or protein representations can be compared with each other. We analyze different types of proteins from the SCOP dataset and show that the Karcher mean and effective dimension of the SRV shape space follow a non-linear pattern as a function of the layers in ESM2 models of different sizes. Furthermore, we use graph filtrations as a tool to study the context lengths at which models encode the structural features of proteins. We find that PLMs preferentially encode immediate as well as local relations between residues, but start to degrade for larger context lengths. The most structurally faithful encoding tends to occur close to, but before the last layer of the models, indicating that training a folding model ontop of these layers might lead to improved folding performance.

蛋白语言模型表征分析结构预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。