提出新型独立性检验方法,无需重抽样即可快速准确判断变量间独立性。
A Martingale Kernel Independence Test
- 用鞅结构构造自标准化统计量,摆脱传统置换校准的高成本
- 在样本量为n时计算复杂度仅为O(n²),且对任意变量数保持高效
- 适合需要高频独立性检测的大规模数据场景,如基因组分析或金融风控
希尔伯特-施密特独立性准则(HSIC)及其联合独立性扩展dHSIC是退化的V统计量,其依赖数据的加权χ²零分布迫使置换校准,使每项测试成本增加两个数量级。本文将近期用于两样本检验的鞅MMD构造推广至(联合)独立性问题,提出两个自标准化统计量,其零分布恒为标准正态,可直接用单次正态分位数查表替代置换步骤。第一个统计量mHSIC是两个经验中心化核矩阵哈达玛积的下三角和,当核函数四阶矩有界且变量独立时收敛于标准正态分布;它对所有固定替代假设具一致性,且在样本量上为二次复杂度,无需样本划分,与有偏的HSIC V统计量相当。第二个统计量mdHSIC通过一次半样本划分实现有限样本一致性:一子样本估计中心化项,另一子样本运行下三角自标准化鞅,使条件均值残差指数级缩小于d,因此在固定变量数下渐近标准正态,每测试开销仅随d线性增长。在输入维度1至500、联合测试2至10个变量的合成数据上,两种统计量的类型Ⅰ误差率与检验功效均匹配置换校准基线,但运行速度提升25至60倍。
原文摘要 · Abstract (English)
The Hilbert-Schmidt Independence Criterion (HSIC) and its joint-independence extension $d\mathrm{HSIC}$ are degenerate $V$-statistics whose data-dependent weighted-$χ^2$ null limits force a permutation calibration that multiplies the per-test cost by the number of permutations, in practice two orders of magnitude. Adapting the recent martingale MMD construction for two-sample testing to the (joint) independence problem, we introduce two studentised statistics whose null distributions are standard normal regardless of the data law, so that a single normal-quantile lookup replaces the permutation step entirely. The first, $m\mathrm{HSIC}$, is a self-normalised lower-triangular sum of the Hadamard product of two empirically centred Gram matrices. Under independence and bounded-fourth-moment kernels it converges to a standard normal. It is consistent against every fixed alternative, and runs at quadratic cost in the sample size without any sample split, matching the biased HSIC $V$-statistic. Our second statistic, $md\mathrm{HSIC}$, achieves finite-sample consistency with a single half-sample split: the centring is estimated on one half and the lower-triangular self-normalised martingale is run on the other, shrinking the conditional-mean residual to a quantity that is exponentially small in $d$, so the statistic is asymptotically standard normal at every fixed number of jointly tested variables, with a per-test cost that grows only linearly in $d$. On synthetic data with per-variable input dimension from $1$ to $500$ and between $2$ and $10$ jointly tested variables, both statistics match the empirical type-I error rate and test power of permutation-calibrated baselines while running $25$ to $60\times$ faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。