arXiv:2603.29529cond-mat.dis-nncs.LG2026-03

用温度采样发现:中温训练能让大模型更好预测蛋白结构。

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

  • 通过朗之万动力学在不同温度下采样损失曲面,分析变压器模型特性。
  • 最优嵌入维度下,中温使多数层参数高度保守,利于学习。
  • 高温和高维嵌入更优预测蛋白接触图,适合结构预测任务。

我们采用统计力学框架,利用朗之万动力学在不同温度下采样基于蛋白质序列数据训练的变压器模型的损失曲面,以刻画低损失流形并理解变压器在蛋白质结构预测中表现优异的机制。与前馈网络不同,变压器的损失不存在一阶相变特征,因此存在一个具有优良学习特性的中温区间。我们发现,当嵌入维度最优时,大部分层的参数在这些温度下高度保守,并提供了确定该维度的有效方法。最后,我们表明,在较高温度和更高嵌入维度下,注意力矩阵对蛋白质接触图的预测能力更强,优于学习最优的条件。

原文摘要 · Abstract (English)

We investigate the parameter space of transformer models trained on protein sequence data using a statistical mechanics framework, sampling the loss landscape at varying temperatures by Langevin dynamics to characterize the low-loss manifold and understand the mechanisms underlying the superior performance of transformers in protein structure prediction. We find that, at variance with feedforward networks, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties. We show that the parameters of most layers are highly conserved at these temperatures if the dimension of the embedding is optimal, and we provide an operative way to find this dimension. Finally, we show that the attention matrix is more predictive of the contact maps of the protein at higher temperatures and for higher dimensions of the embedding than those optimal for learning.

蛋白结构变压器温度采样嵌入维度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。