用核空间表示反事实结果分布,实现更精准的策略评估。
Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings

- 在再生核希尔伯特空间中构建反事实结果分布的嵌入表示
- 提出双重稳健估计器,收敛速度更快且误差更小
- 支持反事实采样和高效假设检验,适合推荐与医疗决策
在推荐、广告和医疗等场景中,估算反事实策略下的结果分布对决策至关重要。本文提出并分析了一种新框架——反事实策略均值嵌入(CPME),将整个反事实结果分布嵌入到再生核希尔伯特空间(RKHS)中,实现灵活且非参数化的分布型离线策略评估。我们引入了插件估计器和双重稳健估计器;后者通过同时校正结果嵌入与倾向性模型的偏差,获得了更优的收敛速率。基于此,我们构建了具有渐近正态性的双重稳健核检验统计量,实现了计算高效的假设检验和置信区间的直接构造。本框架还支持从反事实分布中采样。数值模拟表明,CPME 在实际应用中优于现有方法。
原文摘要 · Abstract (English)
Estimating the distribution of outcomes under counterfactual policies is critical for decision-making in domains such as recommendation, advertising, and healthcare. We propose and analyze a novel framework-Counterfactual Policy Mean Embedding (CPME)-that represents the entire counterfactual outcome distribution in a reproducing kernel Hilbert space (RKHS), enabling flexible and nonparametric distributional off-policy evaluation. We introduce both a plug-in estimator and a doubly robust estimator; the latter enjoys improved convergence rates by correcting for bias in both the outcome embedding and propensity models. Building on this, we develop a doubly robust kernel test statistic for hypothesis testing, which achieves asymptotic normality and thus enables computationally efficient testing and straightforward construction of confidence intervals. Our framework also supports sampling from the counterfactual distribution. Numerical simulations illustrate the practical benefits of CPME over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。