arXiv:2510.11140cs.LG2025-10NeurIPS被引 7

通过显式引入核函数多样性,提升复杂数据的双样本与独立性检验性能。

DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing

  • 设计基于核间协方差的聚合统计量,主动促进核函数多样性。
  • 在多个基准上验证了方法在双样本与独立性测试中的优越性。
  • 适合需要高检测力且处理结构化数据的研究者使用。

为适应复杂结构化数据的核函数双样本与独立性检验,常采用多核聚合以提升检验功效。然而我们发现,直接最大化多核统计量可能导致核函数高度相似、信息重叠,限制聚合效果。为此,我们提出一种显式引入核多样性(基于核间协方差)的聚合统计量。进一步识别出核心挑战:核函数多样性与单个核检验功效之间的权衡——所选核需兼具有效性与差异性。据此提出包含选择推断的检验框架,利用训练阶段信息从学习到的多样核池中筛选个体性能强的核。提供严格的理论证明,涵盖检验功效的一致性与第一类错误控制,并给出所提统计量的渐近分析。最后,通过大量实验证明该方法在多种双样本与独立性测试基准上均表现优异。

原文摘要 · Abstract (English)

To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly maximizing multiple kernel-based statistics may result in highly similar kernels that capture highly overlapping information, limiting the effectiveness of aggregation. To address this, we propose an aggregated statistic that explicitly incorporates kernel diversity based on the covariance between different kernels. Moreover, we identify a fundamental challenge: a trade-off between the diversity among kernels and the test power of individual kernels, i.e., the selected kernels should be both effective and diverse. This motivates a testing framework with selection inference, which leverages information from the training phase to select kernels with strong individual performance from the learned diverse kernel pool. We provide rigorous theoretical statements and proofs to show the consistency on the test power and control of Type-I error, along with asymptotic analysis of the proposed statistics. Lastly, we conducted extensive empirical experiments demonstrating the superior performance of our proposed approach across various benchmarks for both two-sample and independence testing.

核方法假设检验多样性统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。