用函数密度特性判断两变量因果方向,无需假设噪声分布。
Identifying Causal Direction via Dense Functional Classes
- 基于最小描述长度原理,利用具有密度性质的函数类构建因果评分。
- 在真实数据集上,该方法在AUDRC指标上优于现有最佳方法。
- 方法简单可解释,仅一个超参数,适合快速因果推断场景。
本文研究在无隐性混杂因素假设下,两个单变量连续值变量X与Y之间的因果方向判定问题。为区分因果关系,提出一种基于最小描述长度(MDL)原则的双变量因果评分方法,其使用在紧实实区间上具备密度性质的函数类。证明了在特定条件下该评分具有可识别性,且条件易于检验。不假设噪声服从高斯分布,仅要求噪声水平较低。三次样条函数类在紧实区间上具有密度性质,因此被用于实例化该方法,得到名为LCUBE的算法。该方法具备可识别性、可解释性、简单性和高效性,仅有一个超参数。实证评估表明,相较于当前最优方法,LCUBE在真实世界图宾根因果对数据集上实现了更高的AUDRC精度;在10个常见基准数据集上平均精度更优,并在13个数据集上达到高于平均水平的精度。
原文摘要 · Abstract (English)
We address the problem of determining the causal direction between two univariate, continuous-valued variables, X and Y, under the assumption of no hidden confounders. In general, it is not possible to make definitive statements about causality without some assumptions on the underlying model. To distinguish between cause and effect, we propose a bivariate causal score based on the Minimum Description Length (MDL) principle, using functions that possess the density property on a compact real interval. We prove the identifiability of these causal scores under specific conditions. These conditions can be easily tested. Gaussianity of the noise in the causal model equations is not assumed, only that the noise is low. The well-studied class of cubic splines possesses the density property on a compact real interval. We propose LCUBE as an instantiation of the MDL-based causal score utilizing cubic regression splines. LCUBE is an identifiable method that is also interpretable, simple, and very fast. It has only one hyperparameter. Empirical evaluations compared to state-of-the-art methods demonstrate that LCUBE achieves superior precision in terms of AUDRC on the real-world Tuebingen cause-effect pairs dataset. It also shows superior average precision across common 10 benchmark datasets and achieves above average precision on 13 datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。