arXiv:2502.01763cs.LGmath.OC2025-02ICML被引 12

层间预条件法能有效学习特征,理论证明其必要性。

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

  • 用层间预条件法解决梯度下降在非理想输入下的特征学习缺陷。
  • 在线性表示与单指标学习中,标准方法性能明显下降。
  • 适合关注优化器设计与特征学习理论的研究者。

层间预条件法是一类内存高效的优化算法,对每层权重张量的每个轴引入预条件矩阵。这类方法近年来重获关注,在多种神经网络优化任务中表现优于逐元素(“对角”)预条件法,如 Adam(W)。从统计角度出发,我们证明了层间预条件法具有理论必要性。通过分析两个典型模型——线性表示学习与单指标学习,我们发现当输入分布超出理想各向同性假设(即 𝐱∼𝑁(0,𝐼))或病态条件设置时,标准 SGD 会成为次优的特征学习器。我们从理论上和数值上验证了这种次优性是根本性的,而层间预条件法自然地成为解决方案。此外,我们还表明,标准工具如 Adam 预条件法与批归一化仅能轻微缓解此类问题,进一步凸显了层间预条件法的独特优势。

原文摘要 · Abstract (English)

Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods have seen a recent resurgence, demonstrating impressive performance relative to entry-wise ("diagonal") preconditioning methods such as Adam(W) on a wide range of neural network optimization tasks. Complementary to their practical performance, we demonstrate that layer-wise preconditioning methods are provably necessary from a statistical perspective. To showcase this, we consider two prototypical models, linear representation learning and single-index learning, which are widely used to study how typical algorithms efficiently learn useful features to enable generalization. In these problems, we show SGD is a suboptimal feature learner when extending beyond ideal isotropic inputs $\mathbf{x} \sim \mathsf{N}(\mathbf{0}, \mathbf{I})$ and well-conditioned settings typically assumed in prior work. We demonstrate theoretically and numerically that this suboptimality is fundamental, and that layer-wise preconditioning emerges naturally as the solution. We further show that standard tools like Adam preconditioning and batch-norm only mildly mitigate these issues, supporting the unique benefits of layer-wise preconditioning.

优化器特征学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。