研究贝叶斯因果发现如何在隐变量混杂下失效,揭示其后验分布的两种失败模式。
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
- 分析线性高斯模型中两个变量受隐变量混杂时的后验行为
- 发现存在临界相关阈值,超过则错误边被偏好,且样本量越大阈值越低
- 识别出由局部结构决定的两类后验失败模式,适合因果推断研究者参考
贝叶斯因果发现广泛用于通过后验推断量化对有向无环图(DAG)的认知不确定性。然而,其在隐变量混杂下的表现仍不明确,现有研究仅指出混杂破坏可识别性,未刻画后验分布的具体响应。本文针对线性高斯因果模型中两个观测变量间存在加性隐变量混杂的情况,分析后验行为。我们推导出一个关键相关性阈值:当相关性超过该阈值时,评分函数会偏好在混杂变量间引入虚假边;且该阈值随样本量增加而降低——数据越多,越容易偏好错误边。超过此阈值后,我们根据混杂变量周围的局部结构,刻画了两种不同的后验失败模式。研究通过多个图结构的精确后验计算验证了这些预测的失败模式。
原文摘要 · Abstract (English)
Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) through posterior inference. However, its behaviour under latent confounding remains poorly understood, as existing work typically notes that confounding breaks identifiability without characterising how the posterior distribution over DAGs responds. In this work, we analyse posterior behaviour under latent confounding in linear Gaussian causal models, focusing on additive latent confounding between exactly two observed variables. We derive a critical correlation threshold above which the score function favours graphs with a spurious edge between the confounded variables, and show that this threshold decreases with sample size -- more data lowers the correlation required for the spurious edge to be favoured. Beyond this threshold, we characterize two distinct posterior failure regimes determined by the local structure around the confounded variables. Our findings are supported by exact posterior computations on multiple graph structures, demonstrating both the predicted failure regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。