arXiv:2605.04418cs.LGcs.AI2026-05被引 1

揭示大模型预训练中显式约束如何稳定训练并替代传统技巧

Demystifying Manifold Constraints in LLM Pre-training

论文配图:Demystifying Manifold Constraints in LLM Pre-training
图 1 · 摘自论文原文
  • 提出MACRO优化器,通过流形约束控制权重变化
  • 实验证明约束可独立限制激活值尺度并维持稳定旋转平衡
  • 适合研究大模型训练机制或优化算法的学者

大型语言模型(LLM)预训练的成功高度依赖于启发式稳定技术,如显式归一化层和权重衰减。尽管最近的约束优化方法通过显式限制权重可提升数值稳定性与性能,但添加约束的机制与动机仍不清晰。本文系统揭示了显式流形约束在LLM预训练中的作用。通过引入可证明收敛的单循环优化框架——Msign-Aligned Constrained Riemannian Optimizer(MACRO),研究将权重正则化启发式与RMS归一化、解耦权重衰减等相互作用机制分离。理论分析与全面实证表明,流形约束可独立约束前向激活尺度并强制稳定的旋转平衡,从而取代这些启发式机制。在大规模LLM架构上的评估显示,MACRO在保持精确黎曼优化理论保证的同时,实现了极具竞争力的性能。

原文摘要 · Abstract (English)

The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While recent constrained optimization approaches that explicitly restrict weights may improve numerical stability and performance, the mechanism and motivation for adding constraints still remain elusive. This paper systematically demystifies the role of explicit manifold constraints in LLM pre-training. By introducing the Msign-Aligned Constrained Riemannian Optimizer (MACRO)-a provably convergent, single-loop optimization framework-our study disentangles weight regularization heuristics from interacting mechanisms like RMS normalization and decoupled weight decay. Theoretical analyses and comprehensive empirical evaluations reveal that manifold constraints independently bound forward activation scales and enforce stable rotational equilibrium, thereby subsuming the roles of these heuristic mechanisms. Evaluations on large-scale LLM architectures demonstrate that MACRO achieves highly competitive performance while rigorously preserving the theoretical guarantees of exact Riemannian optimization.

大模型训练优化器流形约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。