破解深度模型泛化之谜:从数据依赖的参数空间入手
A Survey on Data-Dependent Worst-Case Generalization Bounds
- 用数据相关的随机假设集扩展PAC-Bayes理论
- 引入分形维数等几何特征细化优化路径复杂度
- 通过稳定性假设替代信息论项,提升实用性
深度神经网络虽参数量巨大却能良好泛化,这与基于固定假设空间的古典学习理论相矛盾。传统在整个参数空间上的统一界在此情境下无意义,近期研究发现若只关注算法实际访问的参数子集,可获得非平凡的泛化保证。本文综述该方向工作,分为三步:将PAC-Bayes理论拓展至数据依赖的随机假设集(arXiv:2404.17442);利用优化轨迹的几何与拓扑描述(如分形维数、α加权寿命和、正幅值)精炼复杂度项(arXiv:2006.09313, arXiv:2302.02766, arXiv:2407.08723);以稳定性假设取代所得信息论项(arXiv:2507.06775)。文章统一这些成果于一个模板不等式,并进行头对头对比分析。
原文摘要 · Abstract (English)
Deep neural networks generalize well despite being heavily overparameterized, in apparent contradiction with classical learning theory based on uniform convergence over fixed hypothesis spaces. Uniform bounds over the entire parameter space are vacuous in this regime, and recent work has shown that non-vacuous guarantees can be recovered by restricting attention to the part of parameter space that the algorithm actually visits. This survey paper organizes this line of work around three steps: extending PAC-Bayesian theory to random, data-dependent hypothesis sets (arXiv:2404.17442); refining the complexity term with geometric and topological descriptors of the optimization trajectory, including fractal dimensions, alpha-weighted lifetime sums, and positive magnitude (arXiv:2006.09313, arXiv:2302.02766, arXiv:2407.08723); and replacing the resulting information-theoretic terms by stability assumptions (arXiv:2507.06775). We unify these contributions around a single template inequality and a head-to-head comparison of the resulting bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。