用观察数据修正偏差,比从零开始设计实验更高效
Observationally Informed Adaptive Causal Experimental Design
- 以观察模型为先验,只优化偏差残差
- 理论证明残差估计收敛更快,节省实验资源
- 适合资源有限但有大量观察数据的场景
随机对照试验(RCT)是因果推断的金标准,但资源稀缺。尽管大规模观察数据普遍存在,却因偏倚担忧仅用于事后分析,未被用于前瞻性试验设计。本文提出主动残差学习新范式,将观察模型作为基础先验,将实验重点从从零学习目标因果量转向高效估计修正观察偏倚所需的残差。为此提出R-Design框架。理论上证明两大优势:(1) 结构性效率差距——估计平滑残差对比具有严格更快的收敛速度;(2) 信息效率——量化了传统参数获取方法(如BALD)在任务无关干扰不确定性上的冗余,浪费预算。提出R-EPIG(残差期望预测信息增益)统一准则,直接针对因果估计算子,最小化残差不确定性或澄清政策决策边界。在合成与半合成基准上实验表明,R-Design显著优于基线,证实修复有偏模型远比从零学习高效。
原文摘要 · Abstract (English)
Randomized Controlled Trials (RCTs) represent the gold standard for causal inference yet remain a scarce resource. While large-scale observational data is often available, it is utilized only for retrospective fusion, and remains discarded in prospective trial design due to bias concerns. We argue this "tabula rasa" data acquisition strategy is fundamentally inefficient. In this work, we propose Active Residual Learning, a new paradigm that leverages the observational model as a foundational prior. This approach shifts the experimental focus from learning target causal quantities from scratch to efficiently estimating the residuals required to correct observational bias. To operationalize this, we introduce the R-Design framework. Theoretically, we establish two key advantages: (1) a structural efficiency gap, proving that estimating smooth residual contrasts admits strictly faster convergence rates than reconstructing full outcomes; and (2) information efficiency, where we quantify the redundancy in standard parameter-based acquisition (e.g., BALD), demonstrating that such baselines waste budget on task-irrelevant nuisance uncertainty. We propose R-EPIG (Residual Expected Predictive Information Gain), a unified criterion that directly targets the causal estimand, minimizing residual uncertainty for estimation or clarifying decision boundaries for policy. Experiments on synthetic and semi-synthetic benchmarks demonstrate that R-Design significantly outperforms baselines, confirming that repairing a biased model is far more efficient than learning one from scratch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。