提出统一框架,用结构化扰动分析提升深度学习泛化界紧致性。
Towards A Unified PAC-Bayesian Framework for Norm-based Generalization Bounds
- 将泛化界推导转化为对异质高斯后验的随机优化问题。
- 通过敏感度矩阵建模参数差异敏感性,获得更紧的泛化界。
- 适用于结构感知与可解释性分析,适合理论研究者参考。
理解深度神经网络的泛化行为仍是现代统计学习理论的核心挑战。现有基于PAC-Bayesian的范数界方法虽具数据依赖性并能捕捉模型的算法与几何特性,但多数依赖各向同性高斯后验、过度使用权重扰动的谱范数集中性,且基本忽略网络架构,限制了界值的紧致性与实际意义。为此,本文提出统一的PAC-Bayesian范数界框架,将泛化界推导重构为对异质高斯后验的随机优化问题。核心是引入敏感度矩阵,量化输出对结构化权重扰动的响应,从而显式纳入参数敏感性的异质性与网络结构信息。在不同结构假设下,该框架导出一族泛化界,可恢复多个已有结果作为特例,并达到或超越当前最优方法的界值。该框架为深度学习提供了一种原则性强、灵活的几何/结构感知与可解释泛化分析方式。
原文摘要 · Abstract (English)
Understanding the generalization behavior of deep neural networks remains a fundamental challenge in modern statistical learning theory. Among existing approaches, PAC-Bayesian norm-based bounds have demonstrated particular promise due to their data-dependent nature and their ability to capture algorithmic and geometric properties of learned models. However, most existing results rely on isotropic Gaussian posteriors, heavy use of spectral-norm concentration for weight perturbations, and largely architecture-agnostic analyses, which together limit both the tightness and practical relevance of the resulting bounds. To address these limitations, in this work, we propose a unified framework for PAC-Bayesian norm-based generalization by reformulating the derivation of generalization bounds as a stochastic optimization problem over anisotropic Gaussian posteriors. The key to our approach is a sensitivity matrix that quantifies the network outputs with respect to structured weight perturbations, enabling the explicit incorporation of heterogeneous parameter sensitivities and architectural structures. By imposing different structural assumptions on this sensitivity matrix, we derive a family of generalization bounds that recover several existing PAC-Bayesian results as special cases, while yielding bounds that are comparable to or tighter than state-of-the-art approaches. Such a unified framework provides a principled and flexible way for geometry-/structure-aware and interpretable generalization analysis in deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。