发现蛋白质模型深度效率低下,多数层贡献有限。
From Words to Amino Acids: Does the Curse of Depth Persist?
- 用探测、扰动等方法分析七类主流蛋白语言模型的层贡献。
- 超过一半的计算任务集中在少数深层,其余层仅微调输出。
- 该现象在序列与结构双模态模型中同样存在,适合模型优化研究者。
蛋白质语言模型(PLMs)已广泛应用于蛋白质工程与从头设计,其架构多采用深度Transformer,通过海量序列数据训练,且通常通过增加模型深度来提升性能。尽管自回归大语言模型(LLMs)揭示了‘深度诅咒’:后期层对最终输出贡献甚微,但这一现象是否存在于非自回归或双模态的PLMs中尚不明确。本文对七类主流PLM家族(涵盖自回归、掩码预测与扩散目标)进行了深度分析,利用统一的探测、扰动及下游评估方法量化各层贡献。结果表明,无论模型规模或训练目标如何,均存在一致的深度依赖模式:任务相关计算高度集中于部分层,其余层仅提供渐进式优化。该趋势在仅序列输入与序列-结构双模态设置中均成立。研究证实,深度低效是现代PLMs的普遍特征,为未来高效架构与训练方法提供了重要启示。
原文摘要 · Abstract (English)
Protein language models (PLMs) have become widely adopted as general-purpose models, demonstrating strong performance in protein engineering and de novo design. Like large language models (LLMs), they are typically trained as deep transformers with next-token or masked-token prediction objectives on massive sequence corpora and are scaled by increasing model depth. Recent work on autoregressive LLMs has identified the Curse of Depth: many later layers contribute little to the final output predictions. These findings naturally raise the question of whether a similar depth inefficiency also appears in PLMs, where many widely used models are not autoregressive, and some are multimodal, accepting both protein sequence and structure as input. In this work, we present a depth analysis of seven popular PLM families across model scales, spanning autoregressive, masked, and diffusion objectives, and quantify how layer contributions evolve with depth using a unified set of probing-, perturbation-, and downstream-evaluation measurements. Across models, we observe consistent depth-dependent patterns that extend prior findings on LLMs: a large fraction of task-relevant computation is concentrated in a subset of layers, while the remaining layers mainly provide incremental refinement of the final prediction. These trends persist beyond sequence-only settings and also appear in multimodal PLMs. Taken together, our results suggest that depth inefficiency is a common feature of modern PLMs, motivating future work on more depth-efficient architectures and training methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。