研究数据分布偏差如何影响基于梯度的因果发现,揭示两类关键偏差及其控制方法。
Effects of Distributional Biases on Gradient-Based Causal Discovery in the Bivariate Categorical Case
- 通过学习边缘或条件分布推断因果结构,是主流梯度因果方法的核心思路。
- 即使在合成数据中,边缘熵不对称和干预偏移不对称也会导致因果推断错误。
- 消除潜在因果因子之间的竞争可提升模型对分布偏差的鲁棒性,适合关注模型可靠性研究者。
基于梯度的因果发现方法在高效、可扩展地从数据中推断因果结构方面展现出巨大潜力,但其易受训练数据分布偏差的影响。本文识别出两类关键偏差:边际分布不对称性(边际熵差异导致因果学习偏向特定因子分解)与边际分布干预偏移不对称性(重复干预使某些变量的分布变化速度更快)。针对具有狄利克雷先验的双变量分类设置,我们展示了这些偏差即便在受控的合成数据中仍可能发生。为评估其影响,我们采用两种简单模型——通过学习边际或条件数据分布来推导因果因子分解,这是梯度因果发现中的常见策略——并验证其对两类偏差的敏感性。此外,我们还展示了如何控制这些偏差。对两种相关现有方法的实证评估表明,消除可能因果因子间的竞争关系可显著提升模型对上述偏差的鲁棒性。
原文摘要 · Abstract (English)
Gradient-based causal discovery shows great potential for deducing causal structure from data in an efficient and scalable way. Those approaches however can be susceptible to distributional biases in the data they are trained on. We identify two such biases: Marginal Distribution Asymmetry, where differences in entropy skew causal learning toward certain factorizations, and Marginal Distribution Shift Asymmetry, where repeated interventions cause faster shifts in some variables than in others. For the bivariate categorical setup with Dirichlet priors, we illustrate how these biases can occur even in controlled synthetic data. To examine their impact on gradient-based methods, we employ two simple models that derive causal factorizations by learning marginal or conditional data distributions - a common strategy in gradient-based causal discovery. We demonstrate how these models can be susceptible to both biases. We additionally show how the biases can be controlled. An empirical evaluation of two related, existing approaches indicates that eliminating competition between possible causal factorizations can make models robust to the presented biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。