arXiv:2606.01184stat.MEcs.AI2026-06

提出新方法捕捉干预导致的分布结构变化,超越平均效应的因果分析。

Topological Ignorability for Structural Causal Effects Beyond Means

  • 用拓扑几何度量(如贝蒂数、欧拉示性数)量化处理组与对照组分布的结构差异
  • 在弱可忽略性不成立时,仍能稳定估计结构效应,而均值效应严重偏差
  • 适用于存在隐藏混杂但结构特征可识别的复杂因果场景,如癌症数据

许多干预并不改变结果分布的均值,而是改变其结构:如将群体分割为孤立子群、产生环或空洞、生成分支或重新组织结果云,却使平均响应几乎不变。此时,基于均值的因果效应(如平均处理效应)可能遗漏关键结构变化。本文引入基于干预后结果分布总结的拓扑-几何因果度量,包括密度超水平贝蒂数、欧拉示性数和持久同调摘要,以量化处理组与对照组在均值之外的结构差异。提出拓扑不可忽略性(topological ignorability),作为条件不可忽略性的拓扑类比,要求所选结构特征不变而非整个反事实分布。当摘要函数为单射时,该条件等价于弱不可忽略性;对于非单射摘要,可在不识别完整干预分布的前提下识别目标结构特征。定义了协变量标准化的拓扑-几何因果效应并开发实用估计器。在两个隐藏混杂基准中验证框架:一个完全合成的精确基准,一个使用威斯康星乳腺癌协变量的真实协变量半合成基准。在两者中,弱不可忽略性均不成立,且平衡协变量几乎消除标准化均值差异,但坐标均值平均处理效应仍存在偏差。相比之下,选定的有限密度超水平贝蒂数和欧拉对比在真实、观察和加权分析中保持稳定。

原文摘要 · Abstract (English)

Many interventions alter the structure of an outcome distribution rather than its mean: they can split a population into disconnected regimes, create loops or holes, generate branches, or reorganize an outcome cloud while leaving the average response nearly unchanged. In such settings, mean-based causal estimands such as the average treatment effect may miss important structural effects. We introduce topological-geometrical causal metrics based on summaries of interventional outcome laws, including density-superlevel Betti summaries, Euler signatures, and persistent-homology summaries. These metrics quantify structural differences between treated and untreated outcome laws beyond averages. We also study the assumptions needed for causal interpretation. We introduce topological ignorability, a topological analogue of conditional ignorability that requires invariance of the chosen structural feature rather than the full counterfactual distribution. When the chosen summary is injective, this condition coincides with weak ignorability; for noninjective summaries, it can identify the structural feature of interest without identifying the full interventional law. We define a covariate-standardized topological-geometrical causal effect and develop practical estimators. We validate the framework in two hidden-confounding benchmarks: a fully synthetic exact benchmark and a real-covariate semi-synthetic benchmark using Wisconsin breast-cancer covariates. In both, weak ignorability fails and balancing observed covariates nearly eliminates standardized mean differences, yet the coordinate-mean average treatment effect remains biased. By contrast, selected finite density-superlevel Betti and Euler contrasts remain stable across oracle, observational, and weighted analyses.

因果推断拓扑学习结构效应隐藏混杂

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。