arXiv:2605.24710cs.LGmath.PR2026-05

揭示宽神经网络特征学习的结构规律,提出可识别与稀疏分解的新理论框架。

Feature Learning in Wide Neural Networks under $μ$P: Identifiability and Sparse-Dictionary Decomposition of the Mean-Field Limit

  • 基于μP参数化,建立带噪声梯度下降的均值场极限唯一性及收敛速率
  • 证明目标函数在特定条件下可实现稀疏字典分解,支持原子数受阈值约束
  • 揭示网络架构与数据匹配的三元学习单元,适用于理论分析与模型设计

在最大更新参数化(μP)下,针对宽两层神经网络,我们建立了四项结构性结果:首先,证明了带噪声梯度下降的均值场极限全局存在且唯一,确定了初始化矩序列允许的最大权重 $w^*$ 为参数矩增长边界,也是流动传播的最大加权矩类;有限粒子近似具有统一时间的平方-Wasserstein 收敛率 $O(N^{-1})$。其次,刻画了均值场极限的可识别性:两个容许参数测度在 $L^2$ 中产生相同网络函数当且仅当其活跃成分在架构的有限秩实现对称性下一致;轨道深度 $D^*_{ ext{orb}}$ 与矩流形深度 $D^*_{ ext{var}}$ 可分离。第三,在 Barron-Hermite 目标条件下,长时间极限测度的活跃支撑可实现稀疏字典分解:其支撑最多包含 $S^*$ 个原子,模有限秩实现对称性,$S^*$ 由显式系数阈值控制。第四,推导出总特征学习误差分解为统计、优化、混沌传播和稀疏残差四部分,目标依赖的 Hermite/Barron 尾部取代初始残差。四项结果通过一个架构恒等式关联:三元组 $(w^*, D^*_{ ext{orb}}, S^*)$——最大允许权重、轨道可识别深度、目标可实现的稀疏字典深度——构成架构-数据对 $(σ, ρ)$ 的自然学习单元。证明除标准 μP 与均值场 Langevin 理论外均为自洽。

原文摘要 · Abstract (English)

We establish four structural results for feature learning in wide two-layer neural networks under the Maximal Update Parametrization ($μ$P). First, we prove global existence and uniqueness of the mean-field limit of noisy gradient descent under $μ$P, identifying the maximal admissible weight $w^*$ on the moment sequence of the initialization as the reciprocal parameter-moment-growth boundary, and hence the largest weighted moment class propagated by the flow. The finite-particle approximation has uniform-in-time squared-Wasserstein rate $O(N^{-1})$. Second, we characterize identifiability of the mean-field limit: two admissible parameter measures induce the same network function in $L^2$ exactly when their active components agree modulo the finite-rank realization symmetry of the architecture. The orbit depth $D^*_{\mathrm{orb}}$ is separated from the moment-variety depth $D^*_{\mathrm{var}}$. Third, under the Barron-Hermite target condition the active support of the long-time limit measure admits a sparse-dictionary decomposition: it is supported on at most $S^*$ atoms modulo finite-rank realization symmetry, with $S^*$ bounded by an explicit coefficient-threshold number. Fourth, we derive the total feature-learning-error decomposition into statistical, optimization, propagation-of-chaos, and sparse-residual components, with a target-dependent Hermite/Barron tail replacing any initialization-only residual. The four results are tied together by an architectural identity: the triple $(w^*, D^*_{\mathrm{orb}}, S^*)$ -- the maximal admissible weight, the orbit identifiability depth, and the sparse-dictionary depth at which the target is realizable -- is the natural learning cell of the architecture-data pair $(σ, ρ)$. The proofs are self-contained except for standard results from $μ$P and mean-field Langevin theory.

神经网络均值场特征学习稀疏分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。