arXiv:2503.21526stat.MLcs.LG2025-03中稿 · the 4th Conference…被引 5

新算法融合分层先验知识,更高效发现带隐变量的因果关系。

Constraint-based causal discovery with tiered background knowledge and latent variables in single or overlapping datasets

  • 基于分层背景知识改进FCI与IOD算法,支持隐变量和重叠数据
  • 新算法在完整使用先验时既正确又完整,效率显著优于旧方法
  • 适合处理复杂现实数据,如医疗、社会科学研究中的不完全观测

本文研究在约束型因果发现中引入分层背景知识。重点在于放宽因果充分性假设,允许因无法测量或未联合采集而存在隐变量,尤其适用于多个重叠数据集场景。我们首先揭示了‘分层FCI’(tFCI)算法的新性质。在此基础上,提出一种结合分层先验知识的IOD算法扩展——‘分层IOD’(tIOD)。结果表明,在充分使用分层背景知识条件下,tFCI与tIOD均具有正确性;而简化版本的tFCI与tIOD兼具正确性与完备性。进一步证明,即使在马尔可夫等价类限制之外,tIOD也通常比IOD更高效且信息量更大。我们给出该性能提升的正式条件。文章附有多个实例,清晰展示分层背景知识的作用与价值。

原文摘要 · Abstract (English)

In this paper we consider the use of tiered background knowledge within constraint based causal discovery. Our focus is on settings relaxing causal sufficiency, i.e. allowing for latent variables which may arise because relevant information could not be measured at all, or not jointly, as in the case of multiple overlapping datasets. We first present novel insights into the properties of the 'tiered FCI' (tFCI) algorithm. Building on this, we introduce a new extension of the IOD (integrating overlapping datasets) algorithm incorporating tiered background knowledge, the 'tiered IOD' (tIOD) algorithm. We show that under full usage of the tiered background knowledge tFCI and tIOD are sound, while simple versions of the tIOD and tFCI are sound and complete. We further show that the tIOD algorithm can often be expected to be considerably more efficient and informative than the IOD algorithm even beyond the obvious restriction of the Markov equivalence classes. We provide a formal result on the conditions for this gain in efficiency and informativeness. Our results are accompanied by a series of examples illustrating the exact role and usefulness of tiered background knowledge.

因果发现隐变量分层知识多数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。