提出深度扩展提升测试性能的理论机制,解释为何增大网络深度能改善模型泛化。
A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks
- 通过分解表示、优化和泛化三方面,分析残差网络深度扩展的改进路径。
- 证明在特定条件下,扩展后模型的总体风险严格低于原模型。
- 适用于研究模型缩放规律或设计更高效深度网络的研究者。
缩放行为——即模型规模和数据量增加时测试性能提升——是现代深度学习的核心经验现象,但其理论基础仍不完整。本文研究归一化残差网络中的深度扩展:从一个训练好的旧假设类出发,在中间层插入新的残差块,探讨何时这种扩展可带来测试风险的可证明降低。我们构建了一个统一框架,将问题分解为表示增益、优化增益和泛化转移。首先,在零初始化附近的梯度下降一阶条件下,证明扩展后的假设类包含一个辅助跳跃模型,其总体风险严格小于原模型。其次,在针对后归一化残差架构定制的范数控制下,建立了扩展模型类的基于范数的Rademacher复杂度界。这些要素导出两种互补的测试风险保证:一种通过总体风险,当存在正总体边界时更紧;另一种直接作用于训练/测试层面,避免霍夫丁转移,对退化情形更具鲁棒性。这些结果共同提供了一个定理驱动的机制,说明在归一化残差网络中深度扩展如何改善测试性能。更广泛地,它们表明缩放本质上是协同的:深度创造新的优化方向,宽度增强弱信号的有限样本可观测性,而数据决定扩展的统计成本是否可控。
原文摘要 · Abstract (English)
The scaling behavior, in which test performance often improves as model size and data increase, is a central empirical phenomenon in modern deep learning, yet its theoretical basis remains incomplete. In this paper, we study depth expansion in normalized residual networks: starting from a trained model in an old hypothesis class, we insert a new residual block at an intermediate layer and ask when such an expansion can yield a provable improvement in test risk. We develop a unified framework that decomposes this question into representational gain, optimization gain, and generalization transfer. First, under a first-order descent condition near zero initialization, we prove that the expanded hypothesis class contains an auxiliary jumpboard model with strictly smaller population risk than the original model. Second, under norm control tailored to post-normalized residual architectures, we establish a norm-based Rademacher complexity bound for the expanded model class. These ingredients lead to two complementary test-risk guarantees: one route passes through population risk and is tighter when a positive population margin is available, while the other works directly at the train/test level, avoids Hoeffding transfer, and is more robust in degenerate regimes. Together, these results provide a theorem-driven mechanism under which residual depth expansion can improve test performance in normalized residual networks. More broadly, they suggest that scaling is inherently joint: depth creates new improving directions, width enhances the finite-sample observability of weak signals, and data determines whether the statistical cost of expansion can be controlled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。