用复杂度度量发现因果关系,提升决策树对因果结构数据的识别能力。
Causal Discovery and Classification Using Lempel-Ziv Complexity
- 基于莱姆佩尔-齐夫复杂度构建因果度量与距离指标
- 在因果生成数据上,因果决策树显著优于传统方法
- 可量化特征因果强度,适合因果推理场景
理解机器学习决策过程中的因果关系是实现可解释人工智能的关键。本文提出一种基于莱姆佩尔-齐夫(Lempel-Ziv, LZ)复杂度的新因果度量及距离度量,并将其应用于决策树,使分裂基于对结果最具因果影响的特征。我们评估了基于因果度量和距离度量的决策树与传统基尼不纯度决策树的性能。尽管整体分类表现相当,但在由因果模型生成的数据集上,因果决策树显著优于其他两种方法。该结果表明,该方法能捕捉经典决策树无法体现的因果结构信息。此外,基于LZ因果度量决策树所选特征,我们提出了每项特征的因果强度,以推断导致结果的主要因果变量。
原文摘要 · Abstract (English)
Inferring causal relationships in the decision-making processes of machine learning algorithms is a crucial step toward achieving explainable Artificial Intelligence (AI). In this research, we introduce a novel causality measure and a distance metric derived from Lempel-Ziv (LZ) complexity. We explore how the proposed causality measure can be used in decision trees by enabling splits based on features that most strongly \textit{cause} the outcome. We further evaluate the effectiveness of the causality-based decision tree and the distance-based decision tree in comparison to a traditional decision tree using Gini impurity. While the proposed methods demonstrate comparable classification performance overall, the causality-based decision tree significantly outperforms both the distance-based decision tree and the Gini-based decision tree on datasets generated from causal models. This result indicates that the proposed approach can capture insights beyond those of classical decision trees, especially in causally structured data. Based on the features used in the LZ causal measure based decision tree, we introduce a causal strength for each features in the dataset so as to infer the predominant causal variables for the occurrence of the outcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。