arXiv:2506.00486cs.LGcs.AI2025-06

发现大模型参数服从广义高斯分布,据此优化训练效率

It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs

  • 基于广义高斯分布设计初始化策略,加速收敛
  • 提出激活与梯度约束训练法,降低冗余与通信开销
  • 适合追求高效分布式训练的AI研发人员

尽管大语言模型(LLMs)发展迅速,其权重、激活值和梯度的统计结构及其对初始化、训练动态和效率的影响仍缺乏深入研究。我们通过实验证明,这些量在LLMs中可被广义高斯(GG)分布良好建模,并据此提出一个统一的端到端优化框架。贡献包括:(1) 一种基于GG分布的初始化方法,与训练后模型统计特性对齐,加速收敛并提升准确率;(2) ACT方法,一种渐进式激活约束训练机制,减少冗余与传播开销;(3) GCT算法,一种梯度约束训练方法,显著降低分布式训练中的通信成本。在多种架构上的实验表明,所提方法可生成更小、更快的模型,且通信开销极低,性能达到或优于标准基线。本工作通过基于统计建模的原理化优化,推动了高效、可扩展、硬件感知的AI系统发展。

原文摘要 · Abstract (English)

Despite rapid progress in large language models (LLMs), the statistical structure of their weights, activations, and gradients-and its implications for initialization, training dynamics, and efficiency-remains largely unexplored. We empirically show that these quantities in LLMs are well modeled by generalized Gaussian (GG) distributions, and introduce a unified, end-to-end optimization framework grounded in this observation. Our contributions are threefold: (1) a GG-based initialization that aligns with trained model statistics, accelerating convergence and improving accuracy; (2) ACT, a progressive activation-constrained training method that reduces redundancy and propagation overhead; and (3) GCT, a gradient-constrained training algorithm that substantially lowers communication cost in distributed training. Experiments across diverse architectures demonstrate consistently smaller, faster models with minimal communication overhead that match or surpass standard baselines. By anchoring LLM optimization in principled statistical modeling, this work advances efficient, scalable, and hardware-aware AI systems.

大模型优化统计建模分布式训练高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。