arXiv:2608.12090cs.LGq-bio.BM2026-08

发现蛋白语言模型中间层比最后一层更适合下游任务。

Task- and dataset-specific information in protein language models

论文配图:Task- and dataset-specific information in protein language models
图 1 · 摘自论文原文
  • 分析13个模型在15个任务中的中间层嵌入表现
  • 浅层嵌入对突变扫描数据表现更好,深层对自然蛋白更优
  • 人工蛋白任务下模型性能显著下降,提示泛化局限

蛋白语言模型(PLMs)将自然语言处理的最新进展引入计算生物学。这些模型在大规模蛋白序列数据上训练,用于将氨基酸序列转换为可用于多种下游任务(DTs)的潜在空间嵌入。通常使用模型最后层的嵌入,但其内部行为仍不清晰。我们分析了13个PLMs在11个数据集上的15个下游任务,研究中间层嵌入的信息量。通过在各层嵌入上训练探测模型并比较性能,评估其潜在空间特征,发现最后层嵌入很少带来最佳下游任务表现。我们还发现任务类型与信息在各层分布之间存在关联:若预训练目标与残基属性预测相关,则信息随层数递增;而全蛋白任务中,数据集特征决定模型表现。包含深度突变扫描(DMS)数据的集合适用浅层嵌入,含多样自然蛋白的数据集则偏好深层嵌入。此外,当任务针对人工蛋白时,模型性能显著下降。

原文摘要 · Abstract (English)

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By a common consensus, embeddings from the model's last layer are used, and the model's internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs from 11 datasets to investigate the informativeness of embeddings created in intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and computed characteristics of the latent spaces they span to estimate the information they contain, and found that the last layers of PLMs rarely contained embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between DTs and the distribution across PLMs' layers of the relevant information to predict that task. For example, similarity between the pre-training objective and the objective of predicting properties of individual residues leads to a steady increase in understanding of such tasks across the layers of PLMs. On the other hand, for whole-protein tasks, we observe that the dataset, rather than the task itself, defines PLMs' ability to perform well on a DT. Embeddings from shallow layers of PLMs perform better for datasets that contain deep mutational scan (DMS) data, while datasets containing diverse natural proteins find most useful embeddings in the models' deeper layers. Additionally, we discover that the performance of PLMs drops significantly when tasks are introduced for artificial proteins.

蛋白语言模型嵌入分析下游任务DMS数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。