解决相关二值数据的因果发现难题,提升真实场景下的因果图学习精度。
Causal Discovery on Dependent Binary Data
- 基于潜变量效用模型与依赖误差,构建去相关方法处理观测间相关性
- 通过配对最大似然估计依赖协方差矩阵,实现潜变量样本去相关化
- 适用于存在单位间依赖的真实数据,如社交网络、生物群体等场景
在因果图学习中,普遍假设数据观测间相互独立,但现实数据常存在单位间的依赖关系,导致结构学习不准确。本文提出一种基于去相关的因果图学习方法,针对依赖的二值数据,采用具有单位间依赖误差的潜变量效用模型定义局部条件分布。通过配对最大似然法估计单位间依赖的协方差矩阵,再利用该矩阵设计一种类似EM的迭代算法,生成并去相关潜变量效用样本,得到可处理的去相关数据。在此基础上,任意标准因果发现方法均可用于学习底层因果图。在合成数据和真实世界数据上的数值实验表明,所提去相关方法显著提升了因果图学习的准确性。
原文摘要 · Abstract (English)
The assumption of independence between observations (units) in a dataset is prevalent across various methodologies for learning causal graphical models. However, this assumption often finds itself in conflict with real-world data, posing challenges to accurate structure learning. We propose a decorrelation-based approach for causal graph learning on dependent binary data, where the local conditional distribution is defined by a latent utility model with dependent errors across units. We develop a pairwise maximum likelihood method to estimate the covariance matrix for the dependence among the units. Then, leveraging the estimated covariance matrix, we develop an EM-like iterative algorithm to generate and decorrelate samples of the latent utility variables, which serve as decorrelated data. Any standard causal discovery method can be applied on the decorrelated data to learn the underlying causal graph. We demonstrate that the proposed decorrelation approach significantly improves the accuracy in causal graph learning, through numerical experiments on both synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。