arXiv:2512.12676stat.MEcs.LG2025-12

用极小化极大中位数聚合提升变分贝叶斯的抗噪能力

Robust Variational Bayes by Min-Max Median Aggregation

  • 将数据分块后,用极小化极大中位数替代均值聚合局部后验
  • 理论证明可实现接近最优的统计收敛速率,且比直接聚合误差更小
  • 适合处理含异常值的大规模数据,尤其适用于有局部隐变量的模型

我们提出一种鲁棒且可扩展的变分贝叶斯(VB)框架,用于有效处理数据集中的污染和异常值。方法将数据划分为 $m$ 个互不重叠的子集,基于鲁棒聚合原则构建联合优化问题。核心洞察是:完整后验分布等价于 $m$ 次幂局部后验分布与均值KL散度最小化器。为增强鲁棒性,用极小化极大中位数形式替代均值KL散度。该设计不仅确保了KL最小化器与证据下界(ELBO)最大化器的一致性,还推动了变分后验均值的统计收敛速率改进。我们发现,在存在局部隐变量时,$m$ 次幂边缘对数似然函数呈现显著差异,因此分别处理两种情形以保障聚合后验的一致性。当存在局部隐变量时,引入‘聚合-缩放’策略。理论上,我们对所提后验进行了非渐近分析,结合对伯恩斯坦-冯·米塞斯(BvM)定理的精细分析,支持 $m$ 随样本量发散的情形。结果表明,两阶段方法相比直接聚合 $m$ 次幂局部后验,近似误差更小。此外,建立了所提后验均值的近乎最优统计速率,推进了极小化极大中位数估计器的理论边界。大量模拟实验验证了方法的有效性。

原文摘要 · Abstract (English)

We propose a robust and scalable variational Bayes (VB) framework designed to effectively handle contamination and outliers in dataset. Our approach partitions the data into $m$ disjoint subsets and formulates a joint optimization problem based on robust aggregation principles. A key insight is that the full posterior distribution is equivalent to the minimizer of the mean Kullback-Leibler (KL) divergence from the $m$-powered local posterior distributions. To enhance robustness, we replace the mean KL divergence with a min-max median formulation. The min-max formulation not only ensures consistency between the KL minimizer and the Evidence Lower Bound (ELBO) maximizer but also facilitates the establishment of improved statistical rates for the mean of variational posterior. We observe a notable discrepancy in the $m$-powered marginal log likelihood function contingent on the presence of local latent variables. To address this, we treat these two scenarios separately to guarantee the consistency of the aggregated variational posterior. Specifically, when local latent variables are present, we introduce an aggregate-and-rescale strategy. Theoretically, we provide a non-asymptotic analysis of our proposed posterior, incorporating a refined analysis of Bernstein-von Mises (BvM) theorem to accommodate a diverging number of subsets $m$. Our findings indicate that the two-stage approach yields a smaller approximation error compared to directly aggregating the $m$-powered local posteriors. Furthermore, we establish a nearly optimal statistical rate for the mean of the proposed posterior, advancing existing theories related to min-max median estimators. The efficacy of our method is demonstrated through extensive simulation studies.

变分贝叶斯鲁棒推断中位数聚合统计理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。