利用大量无标签数据提升两样本检验性能,突破传统方法局限。
A Semi-Supervised Kernel Two-Sample Test

- 基于半监督框架融合协变量信息,构建渐近正态的检验统计量。
- 在固定与局部替代假设下具有一致性,渐近功效显著优于传统核检验。
- 无需复杂校准,适合高维数据中需增强检验力的场景。
我们研究在存在大量未标记协变量数据的半监督设定下的两样本检验问题。标准两样本检验忽略协变量信息,而后者可能显著提升检验性能。然而,引入协变量可能破坏零假设下的可交换性,使校准过程更加复杂。为此,我们提出一种半监督方法,其检验统计量具有渐近正态性,同时有效整合协变量信息。该方法因在零假设下具备渐近正态性,校准简便,并实现远高于无协变量的现有核检验的渐近功效。此外,我们严格证明了该方法在固定与局部替代假设下具有一致性。模拟实验验证了方法在理论和实践上的优势。
原文摘要 · Abstract (English)
We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However, incorporating covariates potentially breaks the exchangeability assumption under the null, which further complicates a calibration procedure. To address these issues, we propose a semi-supervised method that produces a test statistic with asymptotic normality, while effectively integrating additional information from covariates. Our test is straightforward to calibrate due to the asymptotic normality under the null and achieves asymptotic power that is often much higher than existing kernel tests without covariates. Furthermore, we formally show that the proposed method is consistent in power against fixed and local alternatives. Simulations confirm the practical and theoretical strengths of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。