破解交叉熵训练的非凸动态,揭示其收敛至神经坍缩几何的本质。
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
- 通过哈达玛初始化使softmax对角化,将动态简化为奇异值演化。
- 证明梯度流在非凸下全局收敛,即使存在虚假极小点。
- 适用于研究神经坍缩与深度学习优化机制的理论研究者。
交叉熵(CE)训练损失主导深度学习实践,但现有理论常依赖简化假设,如以平方损失替代或限定于凸模型,忽略了关键行为。CE与平方损失产生根本不同的动态,而凸线性模型无法捕捉非凸优化的复杂性。本文深入分析了多类交叉熵优化在非凸区间的动态,研究一个以标准基向量为输入的两层线性神经网络——这是最简非凸扩展,此前其隐式偏差未知。该模型与用于研究神经坍缩的无约束特征模型一致,因此本工作首次证明:在交叉熵上的梯度流会收敛至神经坍缩几何。我们构造了一个显式的李雅普诺夫函数,证明尽管非凸景观中存在虚假临界点,仍可实现全局收敛。核心洞见是:哈达玛初始化使softmax对角化,冻结权重矩阵的奇异向量,使动态完全由奇异值决定。该技术为分析更广泛场景下的交叉熵训练动态开辟了新路径。
原文摘要 · Abstract (English)
Cross-entropy (CE) training loss dominates deep learning practice, yet existing theory often relies on simplifications, either replacing it with squared loss or restricting to convex models, that miss essential behavior. CE and squared loss generate fundamentally different dynamics, and convex linear models cannot capture the complexities of non-convex optimization. We provide an in-depth characterization of multi-class CE optimization dynamics beyond the convex regime by analyzing a canonical two-layer linear neural network with standard-basis vectors as inputs: the simplest non-convex extension for which the implicit bias remained unknown. This model coincides with the unconstrained features model used to study neural collapse, making our work the first to prove that gradient flow on CE converges to the neural collapse geometry. We construct an explicit Lyapunov function that establishes global convergence, despite the presence of spurious critical points in the non-convex landscape. A key insight underlying our analysis is an inconspicuous finding: Hadamard Initialization diagonalizes the softmax operator, freezing the singular vectors of the weight matrices and reducing the dynamics entirely to their singular values. This technique opens a pathway for analyzing CE training dynamics well beyond our specific setting considered here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。