提出新方法生成更全面的因果模型,提升算法评估可靠性。
Addressing pitfalls in implicit unobserved confounding synthesis using explicit block hierarchical ancestral sampling
- 用分块层级祖先采样显式建模未观测混杂因素
- 解决传统方法限制因果图谱和偏相关矩阵的问题
- 适合需要公平评估因果发现算法的研究者
无偏数据合成对评估存在未观测混杂时的因果发现算法至关重要,因真实世界数据集稀缺。常用隐式参数化通过修改异质协方差矩阵非对角线项来编码未观测混杂,同时保持正定性。我们发现现有协议存在两方面问题:其一,对角占优构造限制了偏相关矩阵的谱范围;其二,采样双向边时过度限制图结构,排除了有效因果模型。为此,我们提出改进的显式建模方法,基于分块层级祖先生成真实因果图,并提供将真实DAG转换为祖先图的算法,使因果发现算法输出可对比。我们建立隐式与显式参数化间的联系,证明本方法完全覆盖因果模型空间,包括隐式方法生成的模型,从而实现更稳健的因果发现与推断方法评估。
原文摘要 · Abstract (English)
Unbiased data synthesis is crucial for evaluating causal discovery algorithms in the presence of unobserved confounding, given the scarcity of real-world datasets. A common approach, implicit parameterization, encodes unobserved confounding by modifying the off-diagonal entries of the idiosyncratic covariance matrix while preserving positive definiteness. Within this approach, we identify that state-of-the-art protocols have two distinct issues that hinder unbiased sampling from the complete space of causal models: first, we give a detailed analysis of use of diagonally dominant constructions restricts the spectrum of partial correlation matrices; and second, the restriction of possible graphical structures when sampling bidirected edges, unnecessarily ruling out valid causal models. To address these limitations, we propose an improved explicit modeling approach for unobserved confounding, leveraging block-hierarchical ancestral generation of ground truth causal graphs. Algorithms for converting the ground truth DAG into ancestral graph is provided so that the output of causal discovery algorithms could be compared with. We draw connections between implicit and explicit parameterization, prove that our approach fully covers the space of causal models, including those generated by the implicit parameterization, thus enabling more robust evaluation of methods for causal discovery and inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。