arXiv:2511.16976cs.LGmath.ST2025-11

解析深度平衡模型的梯度下降动态,证明其收敛性与稳定性。

Gradient descent for deep equilibrium single-index models

  • 发现线性DEQ参数在训练中保持在球面上
  • 证明小步长下能线性收敛到全局最优
  • 适用于研究无限深网络训练机制的学者

深度平衡模型(DEQs)最近成为训练无限深共享权重神经网络的强大范式,在多个现代机器学习任务中达到领先性能。尽管实践成功,但对其梯度下降动态的理论理解仍处于活跃研究阶段。本文在简单线性模型和单指标模型设定下,严格研究了DEQ的梯度下降动力学,填补了文献中的若干空白。我们证明线性DEQ存在守恒律,意味着参数在训练过程中始终位于球面上,并利用该性质表明梯度流在所有时间均保持良好条件。进一步证明在合适初始化和足够小步长下,线性DEQ及深度平衡单指标模型的梯度下降可实现线性收敛至全局最小值。最后通过实验验证了理论发现。

原文摘要 · Abstract (English)

Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretically understanding the gradient descent dynamics for training DEQs remains an area of active research. In this work, we rigorously study the gradient descent dynamics for DEQs in the simple setting of linear models and single-index models, filling several gaps in the literature. We prove a conservation law for linear DEQs which implies that the parameters remain trapped on spheres during training and use this property to show that gradient flow remains well-conditioned for all time. We then prove linear convergence of gradient descent to a global minimizer for linear DEQs and deep equilibrium single-index models under appropriate initialization and with a sufficiently small step size. Finally, we validate our theoretical findings through experiments.

深度平衡模型梯度下降收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。