arXiv:2603.04635stat.MLcs.DS2026-03被引 1

利用预测信息提升独立性检验效率,既保可靠又省样本。

Optimal Prediction-Augmented Algorithms for Testing Independence of Distributions

  • 用不靠谱的预测辅助测试,让算法更聪明
  • 预测准时样本量可大幅减少,最坏情况也不出错
  • 适用于高维数据,且达到理论最优效率

独立性检验是统计推断中的基础问题:给定多随机变量联合分布 $p$ 的样本,目标是判断 $p$ 是否为乘积分布,或是否在总变差距离上与所有乘积分布相差至少 $ε$。在非参数有限样本情形下,该任务极耗样本,最小最大样本复杂度随支撑集大小多项式增长。本文突破此最坏情况限制,采用预测增强型分布测试框架,设计可融合辅助但可能不可信预测信息的独立性检验器。该框架确保即使预测错误,测试仍保持最坏情况有效性,而当预测准确时,样本效率显著提升。主要贡献包括:(i) 针对离散二元分布的自适应独立性测试器,其样本复杂度随预测误差动态调整;(ii) 推广至高维多变量场景,用于检验 $d$ 个随机变量的独立性;(iii) 建立匹配的最小最大下界,证明所提测试器达到最优样本复杂度。

原文摘要 · Abstract (English)

Independence testing is a fundamental problem in statistical inference: given samples from a joint distribution $p$ over multiple random variables, the goal is to determine whether $p$ is a product distribution or is $ε$-far from all product distributions in total variation distance. In the non-parametric finite-sample regime, this task is notoriously expensive, as the minimax sample complexity scales polynomially with the support size. In this work, we move beyond these worst-case limitations by leveraging the framework of \textit{augmented distribution testing}. We design independence testers that incorporate auxiliary, but potentially untrustworthy, predictive information. Our framework ensures that the tester remains robust, maintaining worst-case validity regardless of the prediction's quality, while significantly improving sample efficiency when the prediction is accurate. Our main contributions include: (i) a bivariate independence tester for discrete distributions that adaptively reduces sample complexity based on the prediction error; (ii) a generalization to the high-dimensional multivariate setting for testing the independence of $d$ random variables; and (iii) matching minimax lower bounds demonstrating that our testers achieve optimal sample complexity.

独立性检验预测增强最优复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。