梯度下降在张量分解中隐式偏好低管秩解,突破了以往理论局限。
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
- 采用小随机初始化的梯度下降,实现对低管秩解的隐式正则化。
- 首次在非懒惰训练下证明张量分解的隐式正则化效果。
- 适用于图像数据建模,对张量神经网络设计有启发意义。
我们对过参数化的张量分解问题中的隐式正则化现象提供了严格分析,突破了传统懒惰训练范式。针对矩阵分解,该现象已有诸多研究,但通用初始化策略仍难保证梯度下降下的隐式正则化。尽管Cohen等(2016)指出更广泛的神经网络可由张量分解捕捉,但张量情形下的隐式正则化仅在梯度流或懒惰训练下被严格证明。本文首次在梯度下降而非梯度流下建立了张量层面的此类结果。研究聚焦于管状张量积及低管秩概念,该模型在图像数据中具有实际意义。我们证明:在小随机初始化下,过参数化张量分解的梯度下降会表现出向低管秩解的隐式偏倚。数值实验验证了理论预测的动态行为,并凸显小随机初始化的关键作用。
原文摘要 · Abstract (English)
We provide a rigorous analysis of implicit regularization in an overparametrized tensor factorization problem beyond the lazy training regime. For matrix factorization problems, this phenomenon has been studied in a number of works. A particular challenge has been to design universal initialization strategies which provably lead to implicit regularization in gradient-descent methods. At the same time, it has been argued by Cohen et. al. 2016 that more general classes of neural networks can be captured by considering tensor factorizations. However, in the tensor case, implicit regularization has only been rigorously established for gradient flow or in the lazy training regime. In this paper, we prove the first tensor result of its kind for gradient descent rather than gradient flow. We focus on the tubal tensor product and the associated notion of low tubal rank, encouraged by the relevance of this model for image data. We establish that gradient descent in an overparametrized tensor factorization model with a small random initialization exhibits an implicit bias towards solutions of low tubal rank. Our theoretical findings are illustrated in an extensive set of numerical simulations show-casing the dynamics predicted by our theory as well as the crucial role of using a small random initialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。