提出有效对齐维数,解释为何宽模型在测试集上更优。
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension

- 引入可测量的'有效对齐维数',刻画激活梯度的信号噪声结构。
- 证明宽度扩展能降低测试风险,且在有限样本下成立。
- 适用于大模型、残差网络,验证了模型变宽有益的机制。
现有神经网络宽度理论多关注渐近极限,但难以指导从有限训练数据中识别的扩展方向是否在未见数据上依然有效。本文针对保持函数不变的残差扩展,提出有效对齐维数,用于描述激活梯度的信号-噪声几何结构。通过推导独立估计的训练与测试梯度内积的均值和方差,我们得到误对齐概率的有限样本上界。该上界仅依赖有效对齐维数与有效样本量,只需有限二阶矩和非零总体梯度,无需协方差谱假设或预设宽度增长速率。将此证书整合进训练-测试残差扩展框架,获得测试风险改善的高概率条件。在宽度可控的LLaMA风格Transformer、Pythia及ResNet-20上的实验表明,更宽模型具有更大的有效对齐维数和更低的经验误对齐率。直接残差干预验证了该对齐统计量能预测保留损失的变化方向与幅度。
原文摘要 · Abstract (English)
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remains beneficial on unseen data. We study this problem for function-preserving residual expansion and introduce the effective alignment dimension, a measurable quantity describing the signal-noise geometry of activation gradients. By deriving the exact mean and variance of the inner product between independently estimated training and test gradients, we obtain a finite-sample upper bound on misalignment probability. The bound depends only on the effective alignment dimension and an effective sample size, requiring finite second moments and a nonzero population gradient, without covariance spectral assumptions or prescribed width-growth rates. We integrate this certificate into the train-test residual-expansion framework, yielding a high-probability condition for test-risk improvement. Experiments across width-controlled LLaMA-style Transformers, Pythia, and ResNet-20 show that wider models exhibit larger effective alignment dimensions and lower empirical misalignment. Direct residual interventions confirm that the alignment statistic predicts the sign and magnitude of held-out loss changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。