提出六维评估框架,量化因果发现模型在非线性数据中的准确性与可推理性。
Interpretable, multi-dimensional Evaluation Framework for Causal Discovery from observational i.i.d. Data
- 设计距离最优解的六维评估指标,兼顾结构相似性与因果推断能力。
- 首次在七类算法上测试非可识别非线性模式下的性能,揭示近似方法更优。
- 适合关注因果推断可靠性与算法评估的科研人员使用。
从独立同分布的观测数据中进行非线性因果发现需要对生成过程中的结构方程施加严格的可识别性假设。当这些假设被违反时,结构学习方法的评估需采用严谨且可解释的方法,以同时量化估计结果与真实结构的相似性以及所发现图谱用于因果推断的能力。针对现有评估框架的缺失,本文提出一种可解释的六维评价指标——距离最优解(Distance to Optimal Solution, DOS),专为因果发现领域量身定制。此外,本研究首次在七种不同家族的结构学习算法上,基于真实世界过程启发的非可识别非线性因果模式比例递增的场景下进行性能评估。大规模仿真实验包含七个实验因素,结果显示,除基于因果顺序的方法外,近似因果发现方法在逼近最优解方面表现更佳。
原文摘要 · Abstract (English)
Nonlinear causal discovery from observational data imposes strict identifiability assumptions on the formulation of structural equations utilized in the data generating process. The evaluation of structure learning methods under assumption violations requires a rigorous and interpretable approach, which quantifies both the structural similarity of the estimation with the ground truth and the capacity of the discovered graphs to be used for causal inference. Motivated by the lack of unified performance assessment framework, we introduce an interpretable, six-dimensional evaluation metric, i.e., distance to optimal solution (DOS), which is specifically tailored to the field of causal discovery. Furthermore, this is the first research to assess the performance of structure learning algorithms from seven different families on increasing percentage of non-identifiable, nonlinear causal patterns, inspired by real-world processes. Our large-scale simulation study, which incorporates seven experimental factors, shows that besides causal order-based methods, amortized causal discovery delivers results with comparatively high proximity to the optimal solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。