提出ECCIT方法,让条件独立检验更准更可靠。
Empirically Calibrated Conditional Independence Tests
- 用对抗学习找检验偏差,再校正p值
- 小样本和模型错配下仍能控错误率
- 不依赖具体检验方法,适合因果推断
条件独立检验(CIT)广泛用于因果发现和特征选择。即使使用错误发现率(FDR)控制方法,实际中仍常无法提供频数保证。我们指出两种常见失败模式:(i) 小样本时,许多CIT的渐近保证不准确,即使模型正确也难以估计噪声水平和控制误差;(ii) 大样本但模型错配时,未被捕捉的依赖关系会扭曲检验行为,导致零假设下p值不均匀。我们提出经验校准的条件独立检验(ECCIT),通过测量并修正校准偏差。对选定的基线CIT(如GCM、HRT),ECCIT优化一个对抗器,选择特征与响应函数以最大化校准偏差指标,然后拟合单调校准映射,按观测偏差比例调整原检验的p值。在合成与真实数据的多个基准测试中,ECCIT在保持有效FDR的同时,比现有校准策略具有更高检验功效,且为测试无关方法。
原文摘要 · Abstract (English)
Conditional independence tests (CIT) are widely used for causal discovery and feature selection. Even with false discovery rate (FDR) control procedures, they often fail to provide frequentist guarantees in practice. We highlight two common failure modes: (i) in small samples, asymptotic guarantees for many CITs can be inaccurate and even correctly specified models fail to estimate the noise levels and control the error, and (ii) when sample sizes are large but models are misspecified, unaccounted dependencies skew the test's behavior and fail to return uniform p-values under the null. We propose Empirically Calibrated Conditional Independence Tests (ECCIT), a method that measures and corrects for miscalibration. For a chosen base CIT (e.g., GCM, HRT), ECCIT optimizes an adversary that selects features and response functions to maximize a miscalibration metric. ECCIT then fits a monotone calibration map that adjusts the base-test p-values in proportion to the observed miscalibration. Across empirical benchmarks on synthetic and real data, ECCIT achieves valid FDR with higher power than existing calibration strategies while remaining test agnostic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。