arXiv:2607.21806cs.LG2026-07

用反事实正确性约束机器学习决策的因果影响范围

Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness

论文配图:Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness
图 1 · 摘自论文原文
  • 基于反事实正确性与子群体性能的单调性假设构建因果效应边界
  • 利用已有随机试验数据,为新模型部署提供更紧致的因果影响区间
  • 适用于医疗、司法等高风险领域模型更新后的效果评估

预测性机器学习模型正被广泛用于医疗、刑事司法等高风险领域辅助人类决策。评估这些系统对下游结果(如患者生存率或再犯率)的因果影响日益重要。尽管随机对照试验(RCT)能提供高质量证据,但当模型迭代更新时,重复试验往往不可行。本文提出一种部分识别方法,利用历史RCT数据为新模型的因果效应构建边界。核心创新在于引入两个单调性假设:一是个体层面的反事实正确性(在其他条件不变下,正确预测导致不劣于原结果);二是子群体预测性能与结果之间的关系,可解释为对模型输出的信任程度。通过模拟研究验证,该方法可比以往方法得到更紧致的边界。

原文摘要 · Abstract (English)

Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice. There is a growing recognition of the need to evaluate the causal impact of deploying these systems on downstream outcomes, such as patient survival or crime recidivism. Randomized control trials (RCTs) can provide high-quality evidence on the impact of a deployed model, but they run into a challenge: it is often infeasible to run repeated trials when models are updated or retrained to improve predictive performance. In this work, we present a partial-identification approach to using prior RCT data to construct bounds on the causal effect of a new model. The core innovation in our approach is to leverage assumptions relating fine-grained predictive accuracy to downstream outcomes. We do so via two monotonicity assumptions: first, on individual-level `counterfactual correctness' (all else being equal, a correct prediction leads to non-inferior outcomes); and second, on the relation between subgroup predictive performance and outcomes, interpretable as an assumption regarding trust in model outputs. We demonstrate our method with a simulation study, illustrating how incorporating this information can lead to more informative bounds compared to prior work.

因果推断机器学习评估反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。