用数字孪生模拟反事实,让因果推断结果可验证。
The Digital Twin Counterfactual Framework: A Validation Architecture for Simulated Potential Outcomes
- 构建数字孪生仿真器,通过多级验证测试反事实预测
- 将因果效应分为可观测可验和依赖隐藏结构的两类
- 适合需要可验证因果结论的研究者,尤其关注模型可信度
因果推断的根本难题在于:个体的反事实结果永远不可观测。现有方法均依赖假设填补缺失数据,无法直接生成反事实。本文提出数字孪生反事实框架(DTCF):不通过统计估计,而是用数字孪生模拟反事实,并引入分层验证机制。将数字孪生视为潜在结果框架中的随机映射,提出从边际、联合到结构三类孪生保真度假设,逐步释放更丰富的因果估量。核心贡献有三:一是建立五级验证架构,将不可检验的模拟正确性转化为可观测数据上的可检验测试;二是形式化分解因果量,区分可边际验证的(如ATE、CATE、QTE)与依赖联合分布结构的(如ITE分布、获益/受损概率、处理效应方差);三是提供边界、敏感性和不确定性量化工具,显式刻画联合依赖关系。DTCF并未解决因果推断的根本问题,但使边际因果结论越来越可检验,联合因果结论明确依赖假设,且二者间的差距得到形式化刻画。
原文摘要 · Abstract (English)
The fundamental problem of causal inference - that the counterfactual outcome for any individual is never observed - has shaped the entire methodology of the field. Every existing approach substitutes assumptions for missing data: ignorability, parallel trends, exclusion restrictions. None produces the counterfactual itself. This paper proposes the Digital Twin Counterfactual Framework (DTCF): rather than estimating the counterfactual statistically, we simulate it using a digital twin and subject the simulation to a hierarchical validation regime. We formalize the digital twin simulator as a stochastic mapping within the potential outcomes framework and introduce a hierarchy of twin fidelity assumptions - from marginal fidelity through joint fidelity to structural fidelity - each unlocking a progressively richer class of estimands. The central contribution is threefold. First, a five-level validation architecture converts the unfalsifiable claim that the simulator produces correct counterfactuals into falsifiable tests against observable data. Second, a formal decomposition separates causal quantities into those that are marginally validated (ATE, CATE, QTE - testable through observable-arm comparison) and those that are copula-dependent (the ITE distribution, probability of benefit/harm, variance of treatment effects - permanently reliant on the unobservable within-individual dependence structure). Third, bounding, sensitivity, and uncertainty quantification tools make the copula dependence explicit. The DTCF does not resolve the fundamental problem of causal inference. What it provides is a framework in which marginal causal claims become increasingly testable, joint causal claims become explicitly assumption-indexed, and the gap between the two is formally characterized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。