解释为何降低温度能提升生成模型的多样性与真实性。
Understanding temperature tuning in energy-based models
- 基于能量间隙理论,揭示温度调节如何纠正模型对高能态概率的误判。
- 在小样本下,低温采样可显著提升蛋白质序列生成的实用性与多样性。
- 温度调优非一味降低,特定条件下提高温度反而更好,适合模型诊断。
生成复杂系统模型常需后期参数调整以产出有用结果。例如,蛋白质设计中的能量模型通过人为降低采样温度,生成新颖且功能性的序列。这种温度调优是机器学习中常见但理解不足的启发式方法,用于调控生成保真度与多样性之间的权衡。本文提出一个可解释、物理启发的框架来阐明此现象。我们证明,在具有大能量间隙的系统中(即少数有意义状态与海量不现实状态分离),稀疏数据训练会导致模型系统性高估高能态的概率,而降低采样温度可纠正此偏差。更一般地,我们刻画了最优采样温度如何依赖于数据量与系统内在能量景观的相互作用。关键发现是:降低温度并非总是有益;我们识别出在某些条件下提高温度反而能获得更好的生成性能。因此,后验温度调优被重新定义为揭示真实数据分布特性和模型学习局限性的诊断工具。
原文摘要 · Abstract (English)
Generative models of complex systems often require post-hoc parameter adjustments to produce useful outputs. For example, energy-based models for protein design are sampled at an artificially low ''temperature'' to generate novel, functional sequences. This temperature tuning is a common yet poorly understood heuristic used across machine learning contexts to control the trade-off between generative fidelity and diversity. Here, we develop an interpretable, physically motivated framework to explain this phenomenon. We demonstrate that in systems with a large ''energy gap'' - separating a small fraction of meaningful states from a vast space of unrealistic states - learning from sparse data causes models to systematically overestimate high-energy state probabilities, a bias that lowering the sampling temperature corrects. More generally, we characterize how the optimal sampling temperature depends on the interplay between data size and the system's underlying energy landscape. Crucially, our results show that lowering the sampling temperature is not always desirable; we identify the conditions where \emph{raising} it results in better generative performance. Our framework thus casts post-hoc temperature tuning as a diagnostic tool that reveals properties of the true data distribution and the limits of the learned model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。