解决不等样本量下MMD检验的理论难题,无需丢弃数据
Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics
- 基于广义U统计量扩展MMD理论,支持不等样本
- 首次给出非比例情形下的渐近分布,提升检验效力
- 适合实际数据中样本不均衡的研究场景
现有基于最大均值差异(MMD)的两样本检验方法通常假设两个分布的样本量相等。在实际应用中,这可能导致舍弃有价值的数据,降低检验功效。本文通过扩展广义U统计量理论,并将其应用于标准MMD估计器,推导出在不等样本量下(尤其是先前部分结果要求的比例之外)MMD估计器的渐近分布新表征。该推广还提供了一种在不等样本下优化检验功效的新准则。本方法充分利用所有可用数据,提升了检验的准确性和实际适用性。此外,我们给出了更简洁的MMD估计器方差表征,揭示了一个可能令人意外的现象:零MMD意味着估计器退化,但退化估计器也可能对应非零MMD;我们给出了构造与证明,说明在常见情形下此情况不会发生。
原文摘要 · Abstract (English)
Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require discarding valuable data, unnecessarily reducing test power. We address this long-standing limitation by extending the theory of generalized U-statistics and applying it to the usual MMD estimator, resulting in new characterization of the asymptotic distributions of the MMD estimator with unequal sample sizes (particularly outside the proportional regimes required by previous partial results). This generalization also provides a new criterion for optimizing the power of an MMD test with unequal sample sizes. Our approach preserves all available data, enhancing test accuracy and applicability in realistic settings. Along the way, we give much cleaner characterizations of the variance of MMD estimators, revealing something that might be surprising to those in the area: while zero MMD implies a degenerate estimator, it is sometimes possible to have a degenerate estimator with nonzero MMD as well; we give a construction and a proof that it does not happen in common situations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。