用贝叶斯视角解释深度模型过参数化时的双下降现象
Bayesian Double Descent
- 基于贝叶斯框架分析模型复杂度与风险关系
- 证明当参数量超过样本数后风险再次下降
- 适合关注模型泛化与先验设计的研究者
双下降是过参数化统计模型(如深度神经网络)中的一种现象:随着模型复杂度增加,风险先呈U形上升(传统偏差-方差权衡),当参数量等于样本量进入插值区时风险可能无界,随后在超参数化区域再次下降——即双下降。本文揭示该现象具有自然的贝叶斯解释。理论基础包括贝叶斯模型选择、Dickey-Savage密度比,并将广义岭回归与全局-局部收缩方法关联至双下降。我们以高维神经网络为例进行说明,详细分析无限高斯均值模型与非参数回归情形。最后指出未来研究方向。
原文摘要 · Abstract (English)
Double descent is a phenomenon of over-parameterized statistical models such as deep neural networks which have a re-descending property in their risk function. As the complexity of the model increases, risk exhibits a U-shaped region due to the traditional bias-variance trade-off, then as the number of parameters equals the number of observations and the model becomes one of interpolation where the risk can be unbounded and finally, in the over-parameterized region, it re-descends -- the double descent effect. Our goal is to show that this has a natural Bayesian interpretation. We also show that this is not in conflict with the traditional Occam's razor -- simpler models are preferred to complex ones, all else being equal. Our theoretical foundations use Bayesian model selection, the Dickey-Savage density ratio, and connect generalized ridge regression and global-local shrinkage methods with double descent. We illustrate our approach for high dimensional neural networks and provide detailed treatments of infinite Gaussian means models and non-parametric regression. Finally, we conclude with directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。