揭示变分推断中模式崩溃的理论机制,指出即使模型能力足够也会失效。
A theoretical perspective on mode collapse in variational inference
- 通过高斯混合模型分析梯度流,建立低维统计量演化方程
- 发现模式崩溃在理想条件下仍存在,由均值对齐和权重消失驱动
- 结果适用于归一化流等生成模型,为实际应用提供理论依据
深度学习拓展了表达能力强的变分族,但变分推断(VI)的实际效果常受限于传统相对熵目标函数的最小化,可能导致次优解。一个关键挑战是模式崩溃:模型训练时仅聚焦于目标分布的少数模式,尽管其理论上具备表达所有模式的能力。本文对高斯混合模型上的梯度流进行理论研究,识别出刻画流演化的关键低维统计量,并推导出其闭合的低维演化方程。基于该紧凑描述,我们证明即便在统计有利条件下,模式崩溃依然存在,并揭示两种核心驱动力:均值对齐与权重消失。这些理论发现与使用归一化流实现的变分推断一致,提供了实际应用的深刻洞见。
原文摘要 · Abstract (English)
While deep learning has expanded the possibilities for highly expressive variational families, the practical benefits of these tools for variational inference (VI) are often limited by the minimization of the traditional Kullback-Leibler objective, which can yield suboptimal solutions. A major challenge in this context is \emph{mode collapse}: the phenomenon where a model concentrates on a few modes of the target distribution during training, despite being statistically capable of expressing them all. In this work, we carry a theoretical investigation of mode collapse for the gradient flow on Gaussian mixture models. We identify the key low-dimensional statistics characterizing the flow, and derive a closed set of low-dimensional equations governing their evolution. Leveraging this compact description, we show that mode collapse is present even in statistically favorable scenarios, and identify two key mechanisms driving it: mean alignment and vanishing weight. Our theoretical findings are consistent with the implementation of VI using normalizing flows, a class of popular generative models, thereby offering practical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。