用分数生成模型提升条件独立性检验的准确性与稳定性。
Score-based Generative Modeling for Conditional Independence Testing
- 基于分数匹配与朗之万采样,精准建模条件分布。
- 在合成与真实数据集上显著优于现有方法。
- 适合需要高精度因果推断的研究者使用。
确定随机变量间的条件独立关系是机器学习与统计学中的基础但极具挑战性的任务,尤其在高维场景下。现有的基于生成模型的条件独立性检验方法(如利用生成对抗网络)常因条件分布建模不佳和训练不稳定而导致性能欠佳。为此,我们提出一种基于分数生成建模的新方法,实现精确的I类错误控制与强检验功效。具体地,首先采用切片条件分数匹配准确估计条件分数,并利用朗之万动态进行条件采样以生成零假设样本,确保精确的I类错误控制;随后引入拟合优度阶段验证生成样本,提升实际可解释性。我们理论推导了分数生成模型对条件分布建模的误差界,并证明了所提检验的有效性。在合成与真实数据集上的大量实验表明,该方法显著优于现有最先进方法,为生成模型驱动的条件独立性检验提供了有力新途径。
原文摘要 · Abstract (English)
Determining conditional independence (CI) relationships between random variables is a fundamental yet challenging task in machine learning and statistics, especially in high-dimensional settings. Existing generative model-based CI testing methods, such as those utilizing generative adversarial networks (GANs), often struggle with undesirable modeling of conditional distributions and training instability, resulting in subpar performance. To address these issues, we propose a novel CI testing method via score-based generative modeling, which achieves precise Type I error control and strong testing power. Concretely, we first employ a sliced conditional score matching scheme to accurately estimate conditional score and use Langevin dynamics conditional sampling to generate null hypothesis samples, ensuring precise Type I error control. Then, we incorporate a goodness-of-fit stage into the method to verify generated samples and enhance interpretability in practice. We theoretically establish the error bound of conditional distributions modeled by score-based generative models and prove the validity of our CI tests. Extensive experiments on both synthetic and real-world datasets show that our method significantly outperforms existing state-of-the-art methods, providing a promising way to revitalize generative model-based CI testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。