利用损失函数光滑性,为确定性模型推导高概率泛化界。
Smoothness-Based Derandomization of PAC-Bayes Bounds

- 通过后验均值转换吉布斯预测器,量化泛化代价
- 引入参数雅可比与海森矩阵刻画模型平坦度,控制泛化误差
- 适用于线性模型与平滑神经网络,尤其适合批归一化网络
我们研究平滑损失函数下的PAC-Bayes去随机化问题。目标是利用损失和预测器类的光滑性,为确定性预测器获得在高概率下成立的泛化界。我们证明,从吉布斯预测器转移到后验均值处的确定性预测器具有明确的代价,该代价由詹森差距类的泛化间隙决定。我们通过其Rademacher复杂度控制该类,从而得到涉及参数雅可比与得分映射海森矩阵的平坦度量的确定性预测器泛化界。该框架适用于有界和无界平滑损失函数,并特别应用于线性预测器和平滑神经网络。理论中出现的雅可比与海森量启发设计了一个实用正则项。对批归一化网络,我们通过将批归一化变换折叠到相邻仿射权重中,计算其有效权重下的正则项。在CIFAR-10上的实验展示了该正则项在不同批量大小下的行为。
原文摘要 · Abstract (English)
We study PAC-Bayes derandomization for smooth loss functions. Our goal is to obtain generalization bounds that hold with high probability for deterministic predictors by exploiting smoothness properties of both the loss and the predictor class. We show that passing from the Gibbs predictor to the deterministic predictor at the posterior mean has a precise cost, given by the generalization gap of the Jensen gap class. We control this class through its Rademacher complexity, leading to bounds for deterministic predictors that involve flatness quantities expressed in terms of parameter Jacobians and Hessians of the score map. The framework applies to both bounded and unbounded smooth loss functions, and we specialize the results to linear predictors and smooth neural networks. Finally, the Jacobian and Hessian quantities appearing in the theory motivate a practical regularizer. For BatchNorm networks, we compute this regularizer with respect to effective BatchNorm weights obtained by folding the BatchNorm transformation into the adjacent affine weights. Experiments on CIFAR-10 illustrate the behavior of this regularizer under different batch sizes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。