arXiv:2506.01562cs.LGstat.ML2025-06被引 4

调节温度可控制模型表示压缩,提升泛化能力

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization

  • 通过温度调节显式控制软最大值的秩缺陷偏差
  • 低温使表示更压缩,高温提升分布外数据表现
  • 适用于优化分类与注意力模型的表示学习

软最大值函数是深度神经网络的核心组件,广泛用于分类任务的输出分布或Transformer中的注意力权重。尽管应用广泛且效果显著,其对学习动态和学到表示的影响仍不明确,限制了模型行为的优化。本文研究了软最大值在塑造模型表示中的关键作用,提出秩缺陷偏差概念——基于软最大值的深度网络倾向于找到远低于类别数的低秩解。该偏差依赖于软最大值的logits范数,受超参数或软最大值温度直接调控。我们进一步展示如何利用软最大值动态学习压缩表示,或增强其在分布外数据上的性能。在多种架构和真实数据集上验证了结论,表明温度调优具有广泛的适用性,可提升模型表现。本工作为理解软最大值机制提供了新视角,助力更精准地控制深度网络的表示学习。

原文摘要 · Abstract (English)

The softmax function is a fundamental building block of deep neural networks, commonly used to define output distributions in classification tasks or attention weights in transformer architectures. Despite its widespread use and proven effectiveness, its influence on learning dynamics and learned representations remains poorly understood, limiting our ability to optimize model behavior. In this paper, we study the pivotal role of the softmax function in shaping the model's representation. We introduce the concept of rank deficit bias - a phenomenon in which softmax-based deep networks find solutions of rank much lower than the number of classes. This bias depends on the softmax function's logits norm, which is implicitly influenced by hyperparameters or directly modified by softmax temperature. Furthermore, we demonstrate how to exploit the softmax dynamics to learn compressed representations or to enhance their performance on out-of-distribution data. We validate our findings across diverse architectures and real-world datasets, highlighting the broad applicability of temperature tuning in improving model performance. Our work provides new insights into the mechanisms of softmax, enabling better control over representation learning in deep neural networks.

表示学习软最大值温度调节泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。