用信息维数和率失真理论改进因果发现,让算法更准判断因果方向。
Bivariate Causal Discovery Using Rate-Distortion MDL: An Information Dimension Approach
- 基于率失真理论,用信息维数估算原因变量的描述长度。
- 在图宾根数据集上表现优于现有MDL方法,准确率接近最优。
- 适合对因果推断精度要求高的机器学习与统计研究者使用。
基于最小描述长度(MDL)原则的二元因果发现方法,通过近似模型在每个因果方向上的可计算性(不可计算的柯尔莫哥洛夫复杂度),选择总复杂度较低的方向。其前提是真实因果顺序下自然机制更简单。现有方法在估计原因变量的描述长度时存在偏差,实际决策依赖于因果机制的描述长度。本文基于率失真理论,提出一种新方法:以反映底层分布的失真水平为基准,估算原因变量的最小传输速率,该失真水平由基于直方图的密度估计规则推导,而速率则通过信息维数(基于渐近近似)计算。结合传统因果机制描述方式,提出新的二元因果发现方法——率失真MDL(RDMDL)。实验表明,RDMDL在图宾根数据集上表现具有竞争力。所有代码与实验均开源,见github.com/tiagobrogueira/Causal-Discovery-In-Exchangeable-Data。
原文摘要 · Abstract (English)
Approaches to bivariate causal discovery based on the minimum description length (MDL) principle approximate the (uncomputable) Kolmogorov complexity of the models in each causal direction, selecting the one with the lower total complexity. The premise is that nature's mechanisms are simpler in their true causal order. Inherently, the description length (complexity) in each direction includes the description of the cause variable and that of the causal mechanism. In this work, we argue that current state-of-the-art MDL-based methods do not correctly address the problem of estimating the description length of the cause variable, effectively leaving the decision to the description length of the causal mechanism. Based on rate-distortion theory, we propose a new way to measure the description length of the cause, corresponding to the minimum rate required to achieve a distortion level representative of the underlying distribution. This distortion level is deduced using rules from histogram-based density estimation, while the rate is computed using the related concept of information dimension, based on an asymptotic approximation. Combining it with a traditional approach for the causal mechanism, we introduce a new bivariate causal discovery method, termed rate-distortion MDL (RDMDL). We show experimentally that RDMDL achieves competitive performance on the Tübingen dataset. All the code and experiments are publicly available at github.com/tiagobrogueira/Causal-Discovery-In-Exchangeable-Data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。