低秩层可抑制深度网络的泛化误差增长,提升模型泛化能力。
On Generalization Bounds for Neural Networks with Low Rank Layers
- 利用高斯复杂度链式法则分析低秩层对泛化误差的影响
- 证明低秩约束网络的泛化界优于全秩网络
- 为神经坍缩现象提供了新的泛化解释,适合理论研究者
尽管先前优化研究表明深度神经网络倾向于产生低秩权值矩阵,但这种归纳偏置对泛化界的影响仍不明确。本文应用Maurer的高斯复杂度链式法则,分析深度网络中低秩层如何防止秩和维度因子在各层间累积。该方法导出了受秩和谱范数约束网络的泛化界。与以往深度网络的泛化界对比表明,低秩层的深度网络可实现比全秩层更好的泛化性能。此外,该框架也为表现出神经坍缩特性的深度网络的泛化能力提供了新视角。
原文摘要 · Abstract (English)
While previous optimization results have suggested that deep neural networks tend to favour low-rank weight matrices, the implications of this inductive bias on generalization bounds remain underexplored. In this paper, we apply Maurer's chain rule for Gaussian complexity to analyze how low-rank layers in deep networks can prevent the accumulation of rank and dimensionality factors that typically multiply across layers. This approach yields generalization bounds for rank and spectral norm constrained networks. We compare our results to prior generalization bounds for deep networks, highlighting how deep networks with low-rank layers can achieve better generalization than those with full-rank layers. Additionally, we discuss how this framework provides new perspectives on the generalization capabilities of deep networks exhibiting neural collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。