用主动学习高效估计因果效应,减少数据需求。
ActiveCQ: Active Estimation of Causal Quantities
- 基于高斯过程建模回归函数,统一处理多种因果量。
- 通过后验不确定性导出采样策略,显著提升样本效率。
- 适合需要低成本数据的因果推断研究者使用。
估计因果量通常需要大量数据,而获取数据成本高昂,尤其在个体结果测量困难时。为解决现有方法仅关注条件平均处理效应的局限性,本文提出主动估计因果量(ActiveCQ)的通用框架。基于多数因果量可表示为回归函数积分的洞察,框架采用高斯过程建模回归函数,并探索显式密度估计与再生核希尔伯特空间中条件均值嵌入相结合的分布建模方法。后者无需显式密度估计,可在与高斯过程相同函数空间中运作,且随更新自适应优化分布模型。该框架支持从因果量后验不确定性的角度推导采集策略,我们设计了基于信息增益和总方差缩减的两种效用函数。在多种模拟与半合成实验中,所提方法显著优于基线,在各类因果量估计中均实现更高的样本效率。
原文摘要 · Abstract (English)
Estimating causal quantities (CQs) typically requires large datasets, which can be expensive to obtain, especially when measuring individual outcomes is costly. This challenge highlights the importance of sample-efficient active learning strategies. To address the narrow focus of prior work on the conditional average treatment effect, we formalize the broader task of Actively estimating Causal Quantities (ActiveCQ) and propose a unified framework for this general problem. Built upon the insight that many CQs are integrals of regression functions, our framework models the regression function with a Gaussian Process. For the distribution component, we explore both a baseline using explicit density estimators and a more integrated method using conditional mean embeddings in a reproducing kernel Hilbert space. This latter approach offers key advantages: it bypasses explicit density estimation, operates within the same function space as the GP, and adaptively refines the distributional model after each update. Our framework enables the principled derivation of acquisition strategies from the CQ's posterior uncertainty; we instantiate this principle with two utility functions based on information gain and total variance reduction. A range of simulated and semi-synthetic experiments demonstrate that our principled framework significantly outperforms relevant baselines, achieving substantial gains in sample efficiency across a variety of CQs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。