arXiv:2603.00541cs.LGstat.ML2026-03

提出统一谱框架,解决模型宽深缩放下的稳定学习与超参迁移问题。

Spectral Condition for $μ$P under Width-Depth Scaling

  • 基于谱分析建立宽深联合缩放的μP统一框架
  • k≥2时的μP在语言模型中实现稳定特征学习与超参迁移
  • 适用于多种优化器,对Transformer类架构更有效

生成式基础模型正不断增大宽度和深度,给特征学习的稳定性与超参数跨模型规模迁移带来挑战。虽然最大更新参数化(μP)为宽度缩放提供了理论依据,但现有向宽深联合缩放扩展的方法仍零散、依赖架构与优化器,且理论复杂。本文提出一个简洁统一的谱框架,用于μP在宽深联合缩放下的应用。针对含k个变换的深层残差网络,该框架明确了权重及其每步更新范数随宽度和深度的缩放方式。揭示了k=1到k≥2的根本转变,统一了此前分散的μP形式,并指出k≥2更适配包含多分支变换的实用架构(如Transformer)。在此框架基础上,我们通过将谱约束映射为具体超参数设置,推导出适用于广泛优化器的μP通用实现方案,复现并扩展了已有结果。GPT-2风格语言模型实验表明,基于k≥2的μP在宽深缩放下实现稳定特征学习与鲁棒超参迁移,而标准参数化及k=1的μP常失败。结果验证了所提谱框架的有效性。

原文摘要 · Abstract (English)

Generative foundation models are increasingly scaled in both width and depth, posing significant challenges for stable feature learning and reliable hyperparameter (HP) transfer across model sizes. While maximal update parameterization ($μ$P) has provided a principled solution to both problems for width scaling, existing extensions to the joint width-depth scaling regime remain fragmented, architecture- and optimizer-specific, and often rely on technically involved theories. In this work, we develop a simple and unified spectral framework for $μ$P under joint width-depth scaling. For deep residual networks whose residual blocks contain $k$ transformations, the framework specifies how the norms of weights and their per-step updates should scale with width and depth. It reveals a fundamental transition from $k=1$ to $k\geq 2$, unifying previously disparate $μ$P formulations and identifying the $k\geq 2$ case as more appropriate for practical architectures with multi-transformation branches such as Transformers. Building on this framework, we derive a general recipe for implementing $μ$P across a broad class of optimizers by mapping spectral constraints to concrete HP parameterizations, recovering existing results and extending them to additional optimizers. Finally, experiments on GPT-2 style language models show that the $μ$P formulation derived from the $k\geq 2$ case achieves stable feature learning and robust HP transfer under width-depth scaling, whereas standard parameterization and $μ$P in the $k=1$ case often fail to do so. These results support the practical effectiveness of the proposed spectral framework.

参数化模型缩放Transformer谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。