arXiv:2510.07935cs.LGcs.IT2025-10被引 1

提升神经网络泛化风险证书的紧致性,实现更可靠的性能保证。

Some theoretical improvements on the tightness of PAC-Bayes risk certificates for neural networks

  • 基于贝努利分布的KL散度新界,推导出不同经验风险下的最优显式风险界。
  • 提出隐式微分方法,将风险证书优化嵌入模型训练目标函数中。
  • 首次在CIFAR-10上获得非平凡的泛化界,适合关注模型可靠性研究者。

本文提出四项理论改进,提升了基于PAC-Bayes界的风险证书在神经网络中的可用性。首先,推导出两个伯努利分布间KL散度的新界,从而得到不同经验风险范围下分类器真实风险的最紧显式界。其次,提出一种基于隐式微分的高效方法,可将PAC-Bayesian风险证书的优化直接嵌入网络训练的目标函数中。最后,提出一种针对不可导目标(如0-1损失)的边界优化方法。这些理论成果在MNIST和CIFAR-10数据集上进行了实证评估,首次为神经网络在CIFAR-10上提供了非平凡的泛化界。实验代码已公开于github.com/Diegogpcm/pacbayesgradients。

原文摘要 · Abstract (English)

This paper presents four theoretical contributions that improve the usability of risk certificates for neural networks based on PAC-Bayes bounds. First, two bounds on the KL divergence between Bernoulli distributions enable the derivation of the tightest explicit bounds on the true risk of classifiers across different ranges of empirical risk. The paper next focuses on the formalization of an efficient methodology based on implicit differentiation that enables the introduction of the optimization of PAC-Bayesian risk certificates inside the loss/objective function used to fit the network/model. The last contribution is a method to optimize bounds on non-differentiable objectives such as the 0-1 loss. These theoretical contributions are complemented with an empirical evaluation on the MNIST and CIFAR-10 datasets. In fact, this paper presents the first non-vacuous generalization bounds on CIFAR-10 for neural networks. Code to reproduce all experiments is available at github.com/Diegogpcm/pacbayesgradients.

PAC-Bayes泛化界神经网络风险证书

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。