ELBO在特定假设下会过拟合,影响贝叶斯模型选择可靠性。
Occam's Razor is Only as Sharp as Your ELBO
- 通过高斯近似后验的秩假设控制复杂度,影响模型选择结果。
- 低秩假设下ELBO反而导致过拟合,与直觉相反。
- 提醒大规模模型应用中需警惕简化假设对模型选择的误导。
边缘似然(即证据)被视为奥卡姆剃刀的数学体现,可避免过拟合,用于模型选择。变分推断中的证据下界(ELBO)也常用于类似目的。先前研究指出,通过均值场近似限制近似后验族可能导致ELBO欠拟合。本文表明,在一个简单的过参数化回归模型中,基于ELBO的超参数学习也可能产生过拟合,具体取决于高斯近似后验中协方差矩阵的假设秩。令人惊讶的是,在仅存在欠拟合和过拟合两种情形时,证据本身有时会选择过拟合版本,而ELBO则不会。这对希望扩展到大模型的贝叶斯实践者提出警示:为可计算性所作的低秩假设可能影响模型选择能力。
原文摘要 · Abstract (English)
The marginal likelihood, also known as the evidence, is regarded as a mathematical embodiment of Occam's razor, enabling model selection that avoids overfitting. The evidence lower bound (ELBO) objective from variational inference has also been used for similar purposes. Prior work has shown that restricting the approximate posterior family via a mean-field approximation can lead the ELBO to underfit. In this paper, we show how ELBO-based hyperparameter learning in a simple over-parameterized regression model can also produce overfitting, depending on the assumed rank of the covariance matrix in a Gaussian approximate posterior. Surprisingly, among only the underfit and overfit options, Bayesian model selection via the evidence itself sometimes prefers the overfit version, while the ELBO does not. Bayesian practitioners hoping to scale to large models should be cautious about how reduced-rank assumptions needed for tractability may impact the potential for model selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。