提出可处理离散化数据的秩检验方法,提升因果发现准确性
Permutation-Based Rank Test in the Presence of Discretization and Application in Causal Discovery with Mixed Data
- 基于置换法构建新秩检验,适用于混合连续与离散变量
- 在离散化条件下仍能严格控制第一类错误率,优于现有方法
- 适合心理测量等存在有序分类数据的因果推断场景
近年来研究表明,交叉协方差矩阵秩的统计检验在因果发现中具有重要作用,此类检验包含偏相关检验作为特例,并能提供潜变量的额外图结构信息。现有秩检验通常假设所有连续变量均可精确测量,但在实际中许多变量需经过离散化才能观测。例如心理测量中,人格维度的连续水平常被离散为‘不同意’‘中立’‘同意’等有序类别。为此,本文提出混合数据置换秩检验(MPRT),能在部分或全部变量离散化时仍有效控制统计误差。理论上,通过置换方法建立交换性并估计渐近零分布,使得MPRT在离散化情形下仍能有效控制第一类错误;实验上,在合成数据和真实数据上的大量验证表明,该方法在因果发现中兼具有效性与适用性。
原文摘要 · Abstract (English)
Recent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical information about latent variables. Existing rank tests typically assume that all the continuous variables can be perfectly measured, and yet, in practice many variables can only be measured after discretization. For example, in psychometric studies, the continuous level of certain personality dimensions of a person can only be measured after being discretized into order-preserving options such as disagree, neutral, and agree. Motivated by this, we propose Mixed data Permutation-based Rank Test (MPRT), which properly controls the statistical errors even when some or all variables are discretized. Theoretically, we establish the exchangeability and estimate the asymptotic null distribution by permutations; as a consequence, MPRT can effectively control the Type I error in the presence of discretization while previous methods cannot. Empirically, our method is validated by extensive experiments on synthetic data and real-world data to demonstrate its effectiveness as well as applicability in causal discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。