揭示对角线网络在极小初始化下的学习偏置与轨迹规律
Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

- 通过算法等价刻画训练轨迹,推导出优化路径
- 收敛至修正版l1范数最小化解,具稀疏性倾向
- 发现结构不变流形是驱动学习的核心几何机制
本文研究了在极小初始化下,对角线线性网络在回归任务中的梯度流动态。在Pesme & Flammarion (2023)定理1基础上,将分析推广至深层及更广义的两层对角线线性网络(见定义4.1)。具体而言,我们证明这些模型的训练轨迹可被提出的算法1等价刻画,并进一步证明该算法收敛于一个修正的 $ \mathcal{l}_1 $ 范数最小化问题的解。因此,我们建立了两种网络架构在极小初始化下的隐式偏置对应于修正的 $ \mathcal{l}_1 $ 范数。此外,通过识别结构不变流形(SIM)(Zhao et al., 2026)作为主导学习过程的关键几何结构,我们为这些动态背后的机制提供了深入洞察。
原文摘要 · Abstract (English)
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified $ \mathcal{l}_1 $ norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified $ \mathcal{l}_1 $ norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。