针对因果效应估算的采样效率问题,提出直接优化未观测因果量的新主动学习框架。
Causal-EPIG: A Prediction-Oriented Active Learning Framework for CATE Estimation
- 基于因果目标对齐原则,用潜在结果和CATE作为查询价值评估依据。
- 在多个数据集上显著优于传统方法,且最优策略依赖于数据复杂度与基础模型。
- 适合高成本实验中需高效获取处理效应估计的研究者使用。
估计条件平均处理效应(CATE)常受限于结果测量的高成本,主动学习因此至关重要。然而,传统主动学习策略存在根本目标偏差:它们旨在减少模型参数或可观测事实结果的不确定性,而非直接聚焦于真正关注的不可观测因果量。为解决这一错配,我们提出因果目标对齐原则,即采集函数应直接针对不可观测的因果量(如潜在结果和CATE),而非间接代理指标。为此,我们构建了因果EPIG(Causal-EPIG)框架,将信息论中的期望预测信息增益(EPIG)准则扩展至显式量化查询对不可观测因果量不确定性降低的价值。在此统一框架下,我们推导出两种策略:一种全面建模完整因果机制(联合潜在结果),另一种直接聚焦于CATE估计量以实现最高样本效率。大量实验表明,所提策略持续优于标准基线,且关键发现是最佳策略具有上下文依赖性,取决于基础估计器与数据复杂度。该框架为实践中高效进行CATE估计提供了理论指导。
原文摘要 · Abstract (English)
Estimating the Conditional Average Treatment Effect (CATE) is often constrained by the high cost of obtaining outcome measurements, making active learning essential. However, conventional active learning strategies suffer from a fundamental objective mismatch. They are designed to reduce uncertainty in model parameters or in observable factual outcomes, failing to directly target the unobservable causal quantities that are the true objects of interest. To address this misalignment, we introduce the principle of causal objective alignment, which posits that acquisition functions should target unobservable causal quantities, such as the potential outcomes and the CATE, rather than indirect proxies. We operationalize this principle through the Causal-EPIG framework, which adapts the information-theoretic criterion of Expected Predictive Information Gain (EPIG) to explicitly quantify the value of a query in terms of reducing uncertainty about unobservable causal quantities. From this unified framework, we derive two distinct strategies that embody a fundamental trade-off: a comprehensive approach that robustly models the full causal mechanisms via the joint potential outcomes, and a focused approach that directly targets the CATE estimand for maximum sample efficiency. Extensive experiments demonstrate that our strategies consistently outperform standard baselines, and crucially, reveal that the optimal strategy is context-dependent, contingent on the base estimator and data complexity. Our framework thus provides a principled guide for sample-efficient CATE estimation in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。