用有效格拉姆矩阵解释深度网络泛化能力
An Effective Gram Matrix Characterizes Generalization in Deep Networks
- 通过微分方程分析训练中泛化差距演化机制
- 预测测试损失准确,且残差主要分布在小特征值子空间
- 揭示不同数据集与架构下泛化的对齐规律
我们推导出一个微分方程,描述深度网络在梯度下降训练下泛化差距的演化过程。该方程由两个量控制:一个收缩因子使不同数据集对应的轨迹汇聚,一个扰动因子反映在不同数据集上的训练差异。通过分析该方程,我们计算出一个“有效格拉姆矩阵”,其与初始“残差”之间的对齐程度可表征泛化差距。在图像分类数据集上的实证评估表明,该分析能准确预测测试损失。训练过程中,残差主要位于有效格拉姆矩阵最小特征值对应的子空间,说明泛化差距沿训练方向缓慢积累,呈现良性训练过程。本研究通过残差与有效格拉姆矩阵的对齐模式,为不同数据集和架构下的神经网络泛化能力提供了新视角。
原文摘要 · Abstract (English)
We derive a differential equation that governs the evolution of the generalization gap when a deep network is trained by gradient descent. This differential equation is controlled by two quantities, a contraction factor that brings together trajectories corresponding to slightly different datasets, and a perturbation factor that accounts for them training on different datasets. We analyze this differential equation to compute an ``effective Gram matrix'' that characterizes the generalization gap in terms of the alignment between this Gram matrix and a certain initial ``residual''. Empirical evaluations on image classification datasets indicate that this analysis can predict the test loss accurately. Further, during training, the residual predominantly lies in the subspace of the effective Gram matrix with the smallest eigenvalues. This indicates that the generalization gap accumulates slowly along the direction of training, charactering a benign training process. We provide novel perspectives for explaining the generalization ability of neural network training with different datasets and architectures through the alignment pattern of the ``residual" and the ``effective Gram matrix".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。