用谱表示学习提升条件独立性检验的可扩展性与有效性
Toward Scalable and Valid Conditional Independence Testing with Spectral Representations
- 基于偏协方差算子的奇异值分解构造谱表示
- 提出双层对比算法学习表示,理论保证渐近有效性和检验力
- 适合需要高效、可靠因果推断的机器学习研究者
条件独立性(CI)在因果推断、特征选择和图模型中至关重要,但在许多场景下需额外假设才能检验。现有方法常依赖严苛结构假设,限制其有效性。基于核的方法虽更严谨,但适应性与可扩展性不足。本文探索表示学习能否缓解此问题,聚焦于偏协方差算子奇异值分解所得的谱表示,并据此构建简单检验统计量。同时提出一种双层对比算法以学习这些表示。理论证明表示学习误差与检验性能相关,且建立渐近有效性与检验力保证。真实与合成数据实验表明,该方法为可扩展的条件独立性检验提供了一条严谨的统计路径,融合了核方法理论与现代表示学习。
原文摘要 · Abstract (English)
Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural conditions, limiting their validity. Kernel methods using partial covariance operators offer a more principled approach but suffer from limited adaptivity and scalability. In this work, we explore whether representation learning can help address these limitations. Specifically, we focus on representations derived from the singular value decomposition of partial covariance operators and use them to construct a simple test statistic. We also introduce a bi-level contrastive algorithm to learn these representations. Our theory links representation learning error to test performance and establishes asymptotic validity and power guarantees. Experiments on real and synthetic data suggest that this approach offers a principled and statistically grounded path toward scalable CI testing, bridging kernel-based theory with modern representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。