提出统一方差估计方法,提升不平衡数据下MMD检验的准确性与速度。
Unified Unbiased Variance Estimation for Maximum Mean Discrepancy: Robust Finite-Sample Performance with Imbalanced Data and Exact Acceleration under Null and Alternative Hypotheses
- 基于U统计量和Hoeffding分解,统一建模MMD方差在不同假设下的表现
- 在拉普拉斯核下实现一维情况的精确加速,计算复杂度从O(n²)降至O(n log n)
- 适用于样本不均衡场景,适合需要高精度两样本检验的研究者
最大均值差异(MMD)是一种基于核函数的非参数双样本检验统计量,其推断精度高度依赖于方差的准确刻画。现有研究提出了多种有限样本下的MMD方差估计方法,但在零假设与备择假设下、以及平衡或非平衡采样方案中常存在差异。本文通过MMD统计量的U统计量表示及其Hoeffding分解,建立了涵盖不同假设和样本配置的统一有限样本方差表征。基于此分析,我们提出在拉普拉斯核下一维情况的精确加速方法,将整体计算复杂度从 $\mathcal O(n^2)$ 降低至 $\mathcal O(n \log n)$。
原文摘要 · Abstract (English)
The maximum mean discrepancy (MMD) is a kernel-based nonparametric statistic for two-sample testing, whose inferential accuracy depends critically on variance characterization. Existing work provides various finite-sample estimators of the MMD variance, often differing under the null and alternative hypotheses and across balanced or imbalanced sampling schemes. In this paper, we study the variance of the MMD statistic through its U-statistic representation and Hoeffding decomposition, and establish a unified finite-sample characterization covering different hypotheses and sample configurations. Building on this analysis, we propose an exact acceleration method for the univariate case under the Laplacian kernel, which reduces the overall computational complexity from $\mathcal O(n^2)$ to $\mathcal O(n \log n)$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。