arXiv:2608.17434cs.AI2026-08

深度神经网络的泛化误差随深度平方增长,揭示了深层结构的本质限制。

Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression

  • 构建局部打包构造,证明深度平方依赖是内在性质
  • 在特定样本条件下,风险下界达到 L²w²log w 阶
  • 适用于分析深层ReLU网络的泛化能力与表示极限

我们研究在深度L、宽度w、层和变差预算A、输出界B的显式向量值Parhi--Nowak深度-RBV²架构下的高斯回归问题。该架构参数量为O(L w²)。已知上下界相差一个深度因子。通过构造局部打包,证明在显式样本量依赖的半径条件下,深度平方依赖是内在的。该打包的对数基数为Ω(L² w² log w),码字位于O(λ) L²球内,两两间距离至少Ω(λ)。核心工具包括偏差修正的有界系数逼近定理和平衡放大:将深度D的ReLU网络乘以q,可用单个常数通道实现,每系数仅增长q^(1/D)。转换至向量值RBV²块后,层和代价为O(D w² q^(1/D))。高斯Fano给出半径显式下界,受输出、测试和表示尺度影响。当A=B=R,σ∝R且满足所述半径条件时,最小最大风险至少为L²w²log(w) R²/n阶。基于伪维数的有限网上限给出无界高斯响应下O~(L² w² R²/n)。因此,最小最大风险在对数因子范围内呈深度平方多项式依赖,并在小半径下呈现表示受限行为。

原文摘要 · Abstract (English)

We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variation budget A, and output bound B. For this O(L w^2)-parameterized architecture, the known lower and upper bounds differ by one factor of depth. We construct a local packing showing that the quadratic depth dependence is intrinsic under an explicit sample-size-dependent radius condition. The packing has log-cardinality Omega(L^2 w^2 log w); its codewords lie in an O(lambda) L^2 ball and are pairwise Omega(lambda)-separated. The main ingredients are a bias-corrected bounded-coefficient approximation theorem and balanced amplification: multiplying a depth-D ReLU network by q can be implemented using one constant channel so that every coefficient grows by only q^(1/D). Translation to vector-valued RBV^2 blocks then has layer-sum cost O(D w^2 q^(1/D)). Gaussian Fano yields a radius-explicit lower bound governed by the output, testing, and representation scales. Under A=B=R, sigma proportional to R, and the stated radius condition, this gives minimax risk at least of order L^2 w^2 log(w) R^2/n. A pseudodimension-based finite-net upper bound gives O-tilde(L^2 w^2 R^2/n) for unbounded Gaussian responses. Thus the minimax risk has quadratic polynomial dependence on depth, up to logarithmic factors, and exhibits a transition to representation-limited behavior at smaller radius.

深度学习泛化误差神经网络理论统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。