提出一种抗干扰且可扩展的变分贝叶斯方法,适合处理大规模数据中的异常值。
Robust and Scalable Variational Bayes
- 将数据分块计算后验,用沃尔什距离下的几何中位数聚合变分后验。
- 新方法在污染数据下仍保持真实后验的收缩性质,且有理论保障。
- 适用于含复杂噪声的大规模数据,尤其适合对鲁棒性要求高的场景。
我们提出一种稳健且可扩展的变分贝叶斯(VB)框架,能有效处理大规模数据中任意性质的异常值和污染问题。该方法将数据集划分为互不重叠的子集,分别计算各子集的后验,并独立对这些后验应用变分近似。随后,利用沃尔什距离下的概率测度几何中位数对所得变分后验进行聚合,形成新的变分中位后验(VM-Posterior)分布。我们严格证明了该方法在保持与真实后验相似的收缩性质的同时,能考虑变分近似固有的误差或变分间隙。同时提供了VM-Posterior的可证明鲁棒性保证。此外,我们建立了针对具有通用协方差结构的多变量正态分布及均场变分族的变分伯恩斯坦-冯·米塞斯定理。为便于实际应用,我们改进现有算法以计算VM-Posterior,并通过大量数值实验验证其性能。结果表明该方法兼具鲁棒性与可扩展性,是面对复杂污染数据时可靠的贝叶斯推断工具。
原文摘要 · Abstract (English)
We propose a robust and scalable framework for variational Bayes (VB) that effectively handles outliers and contamination of arbitrary nature in large datasets. Our approach divides the dataset into disjoint subsets, computes the posterior for each subset, and applies VB approximation independently to these posteriors. The resulting variational posteriors with respect to the subsets are then aggregated using the geometric median of probability measures, computed with respect to the Wasserstein distance. This novel aggregation method yields the Variational Median Posterior (VM-Posterior) distribution. We rigorously demonstrate that the VM-Posterior preserves contraction properties akin to those of the true posterior, while accounting for approximation errors or the variational gap inherent in VB methods. We also provide provable robustness guarantee of the VM-Posterior. Furthermore, we establish a variational Bernstein-von Mises theorem for both multivariate Gaussian distributions with general covariance structures and the mean-field variational family. To facilitate practical implementation, we adapt existing algorithms for computing the VM-Posterior and evaluate its performance through extensive numerical experiments. The results highlight its robustness and scalability, making it a reliable tool for Bayesian inference in the presence of complex, contaminated datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。