用简单形指数移动平均实现流模型混合的稳定全局加权。
Stable Global Weighting of Flow Mixtures using Simplex Exponential Moving Average

- 通过简单形指数移动平均动态调整多个专家流的混合权重。
- 在10个基准上实现更低的负对数似然,且有效避免组件坍缩。
- 适合需要稳定推理的贝叶斯建模任务,计算开销极低。
归一化流为近似推断提供了强大变分族,但单一架构常难以适应异质后验几何。本文重新审视基于混合的流模型,提出AMF-VI-sEMA框架,采用基于简单形指数移动平均(sEMA)的稳定全局加权机制。第一阶段独立训练多个专家(RealNVP、MAF、RBIG),分别聚焦不同结构模式;第二阶段冻结专家参数,通过温度控制的softmax学习平均对数似然,再在概率单纯形上进行平滑EMA更新。该设计实现无需逐样本门控或反向传播权重的可计算、数据无关的门控机制,自适应重分配容量并防止组件坍缩。在十个后验基准上评估:六个二维合成分布(香蕉形、X形、双峰、多峰、双月、环形)和四个真实/低维贝叶斯目标(BLR、BPR、Weibull、Real-GMM2),对比基线包括NICE、ResFlow和EM-Mixing。评估涵盖负对数似然(NLL)、KL散度、Wasserstein-2距离、最大均值差异(MMD),以及混合动态诊断、超参数敏感性和跨种子鲁棒性。实验表明,AMF-VI-sEMA在所有数据集上均优于前代AMF-VI,避免单一流模型的灾难性传输失败,且保持稳定权重轨迹(所有数据集N_eff > 1.4),计算开销极小。
原文摘要 · Abstract (English)
Normalising flows provide a powerful variational family for approximate inference, yet individual architectures often fail to generalise across heterogeneous posterior geometries. We revisit mixture-based flow formulations and introduce \emph{AMF\mbox{-}VI\mbox{-}sEMA}, a two-stage framework featuring a \emph{stable global weighting} mechanism based on a \emph{Simplex Exponential Moving Average} (sEMA) update. In Stage~1, a heterogeneous set of experts (\textsc{RealNVP}, \textsc{MAF}, \textsc{RBIG}) are trained independently to specialise in distinct structural regimes. In Stage~2, expert parameters are frozen and global mixture weights are learned through a temperature-controlled softmax of average log-likelihoods, followed by a smooth EMA update on the probability simplex. This design produces a tractable, data-agnostic gating mechanism (without per-sample gating or gradient backpropagation through weights) that adaptively reallocates capacity while avoiding component collapse. We evaluate the framework on ten posterior benchmarks: six canonical 2D synthetic families (Banana, X-Shaped, Bimodal, Multimodal, Two-moons, Rings) and four real/low-dimensional Bayesian targets (BLR, BPR, Weibull, Real-GMM2), with stronger baselines (\textsc{NICE}, \textsc{ResFlow}, and EM-Mixing). Comprehensive evaluation covers NLL, KL divergence, Wasserstein-2 distance, and MMD, together with diagnostics of mixture dynamics, hyperparameter sensitivity, and cross-seed robustness. Empirically, \emph{AMF\mbox{-}VI\mbox{-}sEMA} achieves consistent NLL improvements over its predecessor \emph{AMF\mbox{-}VI} and avoids the catastrophic transport failures of single-flow baselines, while maintaining stable weight trajectories ($N_{\mathrm{eff}}{>}1.4$ on all datasets) with minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。