arXiv:2510.26700stat.MLcs.LG2025-10

检验因果机器学习中条件可交换性假设的可靠性,发现违反时模型会误判治疗效果差异。

Assessment of the conditional exchangeability assumption in causal machine learning models: a simulation study

  • 通过模拟实验测试因果森林和X-learner在不同混杂条件下的表现
  • 当条件可交换性被违反时,模型无法正确识别真实异质性,甚至产生虚假异质性
  • 负向控制结果能有效检测子群体混杂,适合用于提升因果推断可信度

观察性研究在构建预测个体化治疗效应(ITEs)的因果机器学习模型时,很少对条件可交换性假设进行实证评估。本文通过模拟研究,考察了因果森林和X-learner模型在存在或不存在真实异质性情况下的混杂偏倚表现。模拟数据涵盖不同混杂程度、样本量及负向控制结果(NCO)的混杂结构。比较了有无未测量混杂时,主结果与NCO上分组治疗效应的估计。当条件可交换性被违反时,两类模型均未能恢复真实治疗效应异质性,某些情况下还错误地显示了本不存在的异质性。尽管NCO未完全满足理想假设,仍能有效识别受未测量混杂影响的子群体,虽不总能精确定位最大混杂子群,但可提示子群体估计中的潜在偏倚。条件可交换性的违反严重削弱了基于常规观察数据的因果机器学习模型对个体化推断的有效性。建议将NCO纳入因果机器学习工作流程,作为检测子群体未测量混杂的实用诊断工具。

原文摘要 · Abstract (English)

Observational studies developing causal machine learning (ML) models for the prediction of individualized treatment effects (ITEs) seldom conduct empirical evaluations to assess the conditional exchangeability assumption. We aimed to evaluate the performance of these models under conditional exchangeability violations and the utility of negative control outcomes (NCOs) as a diagnostic. We conducted a simulation study to examine confounding bias in ITE estimates generated by causal forest and X-learner models under varying conditions, including the presence or absence of true heterogeneity. We simulated data to reflect real-world scenarios with differing levels of confounding, sample size, and NCO confounding structures. We then estimated and compared subgroup-level treatment effects on the primary outcome and NCOs across settings with and without unmeasured confounding. When conditional exchangeability was violated, causal forest and X-learner models failed to recover true treatment effect heterogeneity and, in some cases, falsely indicated heterogeneity when there was none. NCOs successfully identified subgroups affected by unmeasured confounding. Even when NCOs did not perfectly satisfy its ideal assumptions, it remained informative, flagging potential bias in subgroup level estimates, though not always pinpointing the subgroup with the largest confounding. Violations of conditional exchangeability substantially limit the validity of ITE estimates from causal ML models in routinely collected observational data. NCOs serve a useful empirical diagnostic tool for detecting subgroup-specific unmeasured confounding and should be incorporated into causal ML workflows to support the credibility of individualized inference.

因果推断机器学习混杂偏倚负向控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。