arXiv:2603.26024cs.LGcs.LO2026-03

通过条件分布的不对称几何特征,判断两变量间的因果方向。

Identification of Bivariate Causal Directionality Based on Anticipated Asymmetric Geometries

  • 基于预期的条件分布几何不对称性进行因果推断
  • AAG方法在真实数据上达84.3%准确率,优于现有方法
  • 方法对超参数敏感但可通过实验设计优化,适合因果发现任务

二元数值数据中因果方向的识别是基础研究问题,具有重要应用价值。本文提出两种新方法:(1)预期不对称几何(AAG),通过比较实际与预期的条件分布(基于均值和标准差的正态投影);(2)单调性指数(MI),比较沿两轴条件分布梯度的单调性及符号变化次数。两者均假设数据为随机且效应变量的条件分布具有单峰性。采用皮尔逊相关、余弦距离、杰卡德指数、KL散度、K-S距离、MAE、MSE和互信息等多类度量。通过全因子实验设计调优超参数,结果表明AAG方法表现更优,在图宾根真实因果数据集99对样本中,简单调优下准确率达81.4%,自适应调优达84.3%,优于GRCI(81.6%)和CAREFL-H(82.0%)。进一步构建决策树,利用对称统计量区分误判案例,探讨因果识别的确定性程度。

原文摘要 · Abstract (English)

Identification of causal directionality in bivariate numerical data is a fundamental research problem with important practical implications. This paper presents two alternative methods to identify direction of causation by considering conditional distributions: (1) Anticipated Asymmetric Geometries (AAG) and (2) Monotonicity Index (MI). The AAG method compares the actual conditional distributions to anticipated ones along two variables. Different comparison metrics, such as Pearson correlation, cosine distance, Jaccard index, K-L divergence, K-S distance, MAE, MSE, and mutual information have been evaluated. Anticipated distributions have been projected as normal based on dual response statistics: mean and standard deviation. The MI method compares the calculated monotonicity indexes of the gradients of conditional distributions along two axes and exhibits counts of gradient sign changes. Both methods assume stochastic properties of the bivariate data and exploit anticipated unimodality of conditional distributions of the effect. The proposed methods are straightforward and include only a limited number of hyperparameters that affect the accuracy of the identification. For a given set of hyperparameters, both the AAG and MI methods provide a unique, deterministic solution. To address sensitivity to hyperparameters, tuning has been done by utilizing a full factorial Design of Experiment. It turns out that the AAG method outperforms MI, achieving top weighted accuracies of 81.4% with simple tuning and 84.3% with size-adaptive tuning, compared with 81.6% for GRCI or 82.0% for CAREFL-H on the 99 pairs of the Tubingen real-world cause-effect examples. A decision tree has been fitted to distinguish misclassified cases using the input data's symmetrical bivariate statistics to address the question of: How decisive is the identification method of causal directionality?

因果推断二元关系统计建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。