提出新方法检测因果发现中假设违背与小样本误差,无需真实答案即可评估结果可靠性。
How PC-based Methods Err: Towards Better Reporting of Assumption Violations and Small Sample Errors
- 通过分析局部错误传播机制,揭示小误差如何导致全局推断偏差
- 设计一致性评分,在无真值情况下识别更多传统方法遗漏的错误
- 计算开销低,适合作为高成本方法与简单检测间的实用桥梁
基于PC算法的因果发现方法在满足所有结构假设且条件独立性检验正确时具有理论正确性,但现实中这一理想条件极少成立。本文首先分析局部错误如何在PC类方法输出图中传播,揭示看似微小的误差可能造成严重后果。随后提出一致性评分,用于在无真实答案的前提下检测假设违背和小样本误差,该评分仅依赖因果发现算法已执行的统计检验,无需额外测试。本方法检测到的错误类型比现有可比方法更全面。我们将这些计算成本低廉的全局误差检测与量化指标,定位为介于计算昂贵的全局答案集编程方法与较廉价的局部检测方法之间的桥梁。在模拟数据和真实数据集上对评分进行了验证。
原文摘要 · Abstract (English)
Causal discovery methods based on the PC algorithm are proven to be sound if all structural assumptions are fulfilled and all conditional independence tests are correct. This idealized setting is rarely given in real data. In this work, we first analyze how local errors can propagate throughout the output graph of a PC-based method, highlighting how consequential seemingly innocuous errors can become. Next, we introduce coherency scores to find assumption violations and small sample errors in the absence of a ground truth. These scores do not require statistical tests beyond those already executed by the causal discovery algorithm. Errors detected by our approach extend the set of errors that can be detected by comparable existing methods. We place our computationally cheap global error detection and quantification scores as a bridge between computationally expensive global answer-set-programming-based methods and less expensive local error detection methods. The scores are analyzed on simulated and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。