揭示低秩结构对神经网络泛化能力的提升作用
Generalization Bounds for Rank-sparse Neural Networks
- 利用权重矩阵的近似低秩特性分析泛化性能
- 小p时样本复杂度为~O(WrL²),优于传统界
- 适合研究模型泛化与低秩正则的读者
近期研究表明,基于梯度训练的深层神经网络在激活值和权重上呈现瓶颈秩特性:各层激活值的秩趋于一个固定值,即表示训练数据所需的最小秩。这一现象与对线性网络施加权重衰减等价于最小化神经网络的Schatten p quasi范数相一致。本文研究该现象对泛化的影响,证明了若权重矩阵具有近似低秩结构,则可获得更优的泛化界。最终结果依赖于权重矩阵的Schatten p quasi范数:当p较小时,样本复杂度为~O(WrL²),其中W、L分别为网络宽度和深度,r为权重矩阵秩;随着p增大,界趋近于常规范数型界。
原文摘要 · Abstract (English)
It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In fact, the rank of the activations of each layer converges to a fixed value referred to as the ``bottleneck rank'', which is the minimum rank required to represent the training data. This perspective is in line with the observation that regularizing linear networks (without activations) with weight decay is equivalent to minimizing the Schatten $p$ quasi norm of the neural network. In this paper we investigate the implications of this phenomenon for generalization. More specifically, we prove generalization bounds for neural networks which exploit the approximate low rank structure of the weight matrices if present. The final results rely on the Schatten $p$ quasi norms of the weight matrices: for small $p$, the bounds exhibit a sample complexity $ \widetilde{O}(WrL^2)$ where $W$ and $L$ are the width and depth of the neural network respectively and where $r$ is the rank of the weight matrices. As $p$ increases, the bound behaves more like a norm-based bound instead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。