用改进的奈斯特罗姆方法加速大样本两样本检验,兼顾速度与统计可靠性。
A Scalable Nystrom-Based Kernel Two-Sample Test with Permutations
- 基于奈斯特罗姆近似降低最大均值差异(MMD)计算开销
- 在数据分布足够分离时,测试功效达到最优分离率
- 适合大规模科学数据的高效两样本检验,实测验证有效
两样本假设检验——判断两组数据是否来自同一分布——是统计学与机器学习中的基础问题,具有广泛的科学应用。在非参数检验中,最大均值差异(Maximum Mean Discrepancy, MMD)因灵活性和坚实的理论基础而受到青睐。然而,其在大规模场景下的应用受限于高计算成本。本文利用MMD的奈斯特罗姆(Nyström)近似,设计了一种计算高效且实用的检验算法,同时保持统计保证。主要结果是在分布足够分离的条件下,对所提检验的有限样本功效给出了上界,该分离率达到了此设置下的已知极小极大最优率。通过一系列数值实验验证了方法在真实科学数据上的适用性。
原文摘要 · Abstract (English)
Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing, maximum mean discrepancy (MMD) has gained popularity as a test statistic due to its flexibility and strong theoretical foundations. However, its use in large-scale scenarios is plagued by high computational costs. In this work, we use a Nyström approximation of the MMD to design a computationally efficient and practical testing algorithm while preserving statistical guarantees. Our main result is a finite-sample bound on the power of the proposed test for distributions that are sufficiently separated with respect to the MMD. The derived separation rate matches the known minimax optimal rate in this setting. We support our findings with a series of numerical experiments, emphasizing applicability to realistic scientific data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。