破解深度高斯过程变分推断中的后验坍缩问题,提出无需约束先验的初始化方法。
An Analysis of Posterior Collapse, Parameterization and Initialization in Variational Deep Gaussian Processes

- 发现线性先验均值改善优化条件,而非避免深层病态
- 提出零均值先验的新初始化,可避免后验坍缩
- 白化参数化提升收敛稳定性,利于模型训练
深度高斯过程(DGPs)是多层组合高斯过程的生成式概率模型,具备优异预测性能。但由于精确推断不可行,通常采用变分推断(VI)近似后验分布,通过最小化KL散度进行优化。然而,变分推断常面临后验坍缩问题,即变分后验退化为先验,导致数据被解释为噪声。本文分析了该现象在变分深度高斯过程中的成因,揭示其与DSVI算法及前几层普遍使用的线性先验均值函数密切相关。研究发现,线性先验均值的优势并非源于避免深层非单射病态,而是改善了初始优化条件。为此,提出一种零先验均值的替代初始化方案,使模型在初始化时表现如同具有线性先验均值。该方法允许基于建模假设选择先验,而非受优化便利性限制。分析涵盖三种常见参数化形式,表明并非所有参数化都受益于线性先验均值。此外,解释了白化参数化为何能实现更稳定的收敛——这一经验性观察首次获得理论支持,且其稳定性有助于缓解后验坍缩。大量实验验证:所提初始化可有效防止后验坍缩、提升训练稳定性,并达到甚至超越线性先验均值的性能。
原文摘要 · Abstract (English)
DGPs are probabilistic models with remarkable prediction performance that concatenate GPs across several layers. Exact inference in DGPs is intractable, and variational inference is often used to approximate the posterior with a parametric distribution tuned by minimizing the Kullback-Leibler divergence. Moreover, finding a good VI approximation is challenging. In particular, a problem of VI is posterior collapse, where VI converges to a variational posterior that matches the prior. In variational DGPs, this implies explaining the data as noise. This work studies posterior collapse in DGPs and identifies its connection to the DSVI algorithm and the widely used linear prior mean function employed in all but the last layer. We show that the benefit of the linear prior mean does not arise from avoiding the non-injective pathology in very deep DGPs, as previously believed, but from improving the conditioning of the optimization problem at initialization. Thus, we propose an alternative initialization of a zero prior mean DGP that mimics a DGP with a linear prior mean at initialization. This enables successful training of DGPs without imposing optimization-driven constraints on the prior, allowing to choose the prior based on modeling assumptions rather than optimization convenience. Our analysis considers three common parameterizations of DGPs and shows that not all of them benefit from a linear prior mean. We also explain why a whitened parameterization of the \DGP provides more stable convergence, something often assumed from experience, but lacking a rigorous analysis. Furthermore, we show that this stability is also beneficial to avoid the posterior collapse problem. Extensive experiments validate our findings: the proposed initialization prevents posterior collapse, improves stability, and achieves performance comparable to (and sometimes better than) DGPs with a linear prior mean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。