用特征谱复杂度解释深度网络为何泛化,比传统方法更精准。
Pointwise Generalization in Deep Neural Networks

- 基于各层特征的谱信息定义点态黎曼维数,刻画模型复杂度。
- 理论与实验均显示泛化界比传统方法紧一个数量级以上。
- 适合研究模型泛化机制或优化隐性偏好的学者参考。
我们针对深度神经网络为何能泛化这一根本问题,建立了全连接网络的点态泛化理论。该框架突破了长期难以刻画非线性特征学习区间的瓶颈,为表示学习构建了新的统计基础。对每个训练好的模型,我们通过各层学习到的特征表示的特征值,推导出其点态黎曼维数,从而建立依赖于假设、感知表示的泛化边界。这些边界系统性优于基于模型大小、范数乘积及无限宽线性化的方法,在理论和实验中均实现数量级上的改进。理论上,我们揭示了深度网络可计算性的结构特性和数学原理;实验上,点态黎曼维数表现出显著的特征压缩,随过参数化程度增加而下降,并捕捉优化器的隐式偏好。综合来看,结果表明深度网络在实际条件下具有数学可处理性,其泛化能力可通过点态、特征谱感知的复杂度精确解释。
原文摘要 · Abstract (English)
We address the fundamental question of why deep neural networks generalize by establishing a pointwise generalization theory for fully connected networks. This framework resolves long-standing barriers to characterizing the rich nonlinear feature-learning regime and builds a new statistical foundation for representation learning. For each trained model, we characterize the hypothesis via a pointwise Riemannian Dimension, derived from the eigenvalues of the learned feature representations across layers. This establishes a principled framework for deriving hypothesis-dependent, representation-aware generalization bounds. These bounds offer a systematic upgrade over approaches based on model size, products of norms, and infinite-width linearizations, yielding guarantees that are orders of magnitude tighter in both theory and experiment. Analytically, we identify the structural properties and mathematical principles that explain the tractability of deep networks. Empirically, the pointwise Riemannian Dimension exhibits substantial feature compression, decreases with increased over-parameterization, and captures the implicit bias of optimizers. Taken together, our results indicate that deep networks are mathematically tractable in practical regimes and that their generalization is sharply explained by pointwise, feature-spectrum-aware complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。