提出梯度传输新框架,揭示大模型训练中规模与效率的内在关系
Finite-Size Gradient Transport in Large Language Model Pretraining: From Cascade Size to Intensive Transport Efficiency

- 基于五个可观测量构建有限尺寸梯度传输框架,分离规模与效率
- 发现不同模型在传输效率上表现迥异,但共享近似单位规模基线
- 适用于分析训练过程中的动态特性,适合研究模型缩放规律的学者
我们提出一个用于真实语言模型训练的有限尺寸梯度传输框架,基于五个可观测量(D, z, β, δ, v_rel),可分离级联规模、持续时间、绝对传输量及强度传输效率。通过对四尺度125个对齐步长的Pico-LM原始梯度数据,以及由153个对齐检查点差值更新场构建的五尺度Pythia配套数据集进行分析,两种体系均满足相同的代数闭合关系,且均具有接近单位的级联规模主干。但二者处于不同的传输状态:Pico-LM呈现正持续时间标度和负强度效率标度,而Pythia则保持在D=1基准附近,仅具微弱的正效率标度依赖。随机场对照实验在强度与持续时间通道中给出几乎一致的零基线,表明差异源于对共同零基线的真实偏离,而非不同校准。两组在逐步幂律压缩性上也存在差异:Pico-LM保持清晰的持续时间与效率幂律,而Pythia虽维持规模主干,但在这些通道中压缩性较弱。外部性能关联表现为通道层面,主要由v_rel与归一化级联持续时间驱动,而D(t)作为共享规模主干,未表现出显著的指数级性能关联。结果支持一种可复用的传输测量框架,但不主张存在普适固定点或神经缩放律的首因推导。
原文摘要 · Abstract (English)
We introduce a finite-size gradient-transport framework for real language-model training, based on five observables $(D,z,β,δ,v_{\mathrm{rel}})$ that separate cascade size, duration, absolute transport, and intensive transport efficiency. We analyze direct raw-gradient measurements from Pico-LM across four scales and 125 aligned steps, together with a five-scale Pythia companion dataset built from 153 aligned checkpoint-difference update fields. The same algebraic closure holds in both families, and both share a near-unity cascade-size backbone, but they occupy distinct transport regimes: Pico-LM shows positive duration scaling and negative intensive-efficiency scaling, whereas Pythia remains near the $D=1$ baseline with only weak positive efficiency scale dependence. Randomized-field controls give nearly matched null floors in the intensive and duration channels, indicating that the contrast reflects different real departures from a shared null skeleton rather than different null calibrations. The families also differ in stepwise power-law compressibility: Pico-LM retains clean duration and efficiency power laws, whereas Pythia preserves the size backbone but shows weaker one-slope compressibility in those channels. External performance associations are correspondingly channel-level, carried mainly by $v_{\mathrm{rel}}$ and normalized cascade duration, while $D(t)$ acts as a shared size backbone without a significant exponent-level performance association. These results support a reusable transport measurement framework without claiming a universal fixed point or a first-principles derivation of neural scaling laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。