用可解释树结构识别数据中的悖论效应,帮非专家做出更可靠决策。
De-paradox Tree: Breaking Down Simpson's Paradox via A Kernel-Based Partition Algorithm
- 基于核方法递归划分数据,自动发现隐藏子群体
- 在假设因果关系下,准确识别嵌套的相反效应
- 适合缺乏因果推断背景的从业者使用
真实世界的观察数据与机器学习推动了数据驱动决策的发展,但许多模型依赖于可能受混杂因素和子群体异质性影响的实证关联。辛普森悖论正是此类挑战的典型表现:聚合层面与子群体层面的关联方向相反,导致误导性结论。现有方法在检测和解释此类悖论关联方面支持有限,尤其对无深度因果知识的实践者而言。我们提出 De-paradox Tree,一种可解释算法,旨在揭示在假定包含混杂因子和效应异质性的因果结构下,悖论关联背后的隐藏子群体模式。该算法采用新颖的分裂准则与平衡化过程,通过递归划分调整混杂因素并同质化异质效应。相比前沿方法,De-paradox Tree 构建更简洁、更可解释的树结构,能选择相关协变量,并识别嵌套的相反效应,同时在提供因果可接受变量时确保因果效应的稳健估计。本方法通过引入可解释框架,弥补传统因果推断与机器学习方法的局限,支持非专家实践者,明确标注因果假设与适用范围,提升复杂观察数据环境下的决策可靠性。
原文摘要 · Abstract (English)
Real-world observational datasets and machine learning have revolutionized data-driven decision-making, yet many models rely on empirical associations that may be misleading due to confounding and subgroup heterogeneity. Simpson's paradox exemplifies this challenge, where aggregated and subgroup-level associations contradict each other, leading to misleading conclusions. Existing methods provide limited support for detecting and interpreting such paradoxical associations, especially for practitioners without deep causal expertise. We introduce De-paradox Tree, an interpretable algorithm designed to uncover hidden subgroup patterns behind paradoxical associations under assumed causal structures involving confounders and effect heterogeneity. It employs novel split criteria and balancing-based procedures to adjust for confounders and homogenize heterogeneous effects through recursive partitioning. Compared to state-of-the-art methods, De-paradox Tree builds simpler, more interpretable trees, selects relevant covariates, and identifies nested opposite effects while ensuring robust estimation of causal effects when causally admissible variables are provided. Our approach addresses the limitations of traditional causal inference and machine learning methods by introducing an interpretable framework that supports non-expert practitioners while explicitly acknowledging causal assumptions and scope limitations, enabling more reliable and informed decision-making in complex observational data environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。