提出秩约束方法,解决选择偏差下的隐变量因果发现难题。
Latent Variable Causal Discovery under Selection Bias
- 用协方差子矩阵的秩约束替代传统条件独立性,捕捉选择偏差影响。
- 在选择偏差下仍可识别一因子模型的因果结构与选择机制。
- 适用于存在隐藏变量且数据有选择偏差的因果推断场景。
解决隐变量因果发现中的选择偏差问题至关重要但研究不足,主要因缺乏合适的统计工具:尽管已有多种超越基本条件独立性的工具用于处理隐变量,但尚未有适用于选择偏差的方法。本文通过研究秩约束展开尝试,该约束作为条件独立性的推广,利用线性高斯模型中协方差子矩阵的秩。我们证明,尽管选择会显著改变联合分布,但偏倚协方差矩阵中的秩仍保留关于因果结构和选择机制的有意义信息。本文给出了此类秩约束的图论刻画,并据此证明在一因子模型下,即使存在选择偏差,因果结构仍可被识别。仿真与真实数据实验验证了所提秩约束的有效性。
原文摘要 · Abstract (English)
Addressing selection bias in latent variable causal discovery is important yet underexplored, largely due to a lack of suitable statistical tools: While various tools beyond basic conditional independencies have been developed to handle latent variables, none have been adapted for selection bias. We make an attempt by studying rank constraints, which, as a generalization to conditional independence constraints, exploits the ranks of covariance submatrices in linear Gaussian models. We show that although selection can significantly complicate the joint distribution, interestingly, the ranks in the biased covariance matrices still preserve meaningful information about both causal structures and selection mechanisms. We provide a graph-theoretic characterization of such rank constraints. Using this tool, we demonstrate that the one-factor model, a classical latent variable model, can be identified under selection bias. Simulations and real-world experiments confirm the effectiveness of using our rank constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。