基于分布结果的离线策略学习,用瓦斯歇斯坦均值优化个体化决策。
Wasserstein Policy Learning for Distributional Outcomes
- 用瓦斯歇斯坦均值定义多维结果的效用,构建新型策略学习框架。
- 理论证明有限样本后悔率与策略类复杂度和样本数呈平方根关系。
- 适用于因果推断中处理结果为分布的场景,如医疗或金融风险评估。
离线策略学习在因果推断中日益受到关注,其目标是学习一个从协变量到处理措施的映射(个性化处理规则),以最大化基于标量潜在结果均值定义的经验福利。本文研究具有分布值结果的离线策略学习,其中每个潜在结果是ℝ上的概率测度,奖励通过作用于诱导结果分布的瓦斯歇斯坦均值的效用泛函定义。我们基于逆概率加权(IPW)和双重稳健(DR)估计器,建立了该策略学习框架的统计保证。通过处理策略类与无限维分位数域乘积上的均匀偏差挑战,证明了有限样本后悔率的主导依赖关系为~𝒪(√(N-dim(Π)/N))。在一维瓦斯歇斯坦设置下,在给定正则性条件下,主导后悔率仍由策略类复杂度决定。此外,我们提供了极小极大下界,证明了对N和N-dim(Π)的主导依赖关系的紧致性。
原文摘要 · Abstract (English)
Offline policy learning has received growing attention in causal inference. The primary objective is to learn a policy (individualized treatment rule) as a mapping from covariates to treatment that maximizes the empirical welfare defined as the mean of scalar-valued potential outcomes. In this paper, we study offline policy learning with distribution-valued outcomes, where each potential outcome is a probability measure on $\mathbb{R}$ and the reward is defined through a utility functional applied to the Wasserstein barycenter of induced outcome distributions. We establish statistical guarantees for the policy learning framework based on both Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators. By handling the challenging uniform deviation over the product of the combinatorial policy class and the infinite-dimensional quantile domain, we prove that the finite-sample regret has leading dependence $\widetilde{\mathcal{O}}(\sqrt{\mathrm{N\text{-}dim}(Π)/N})$. In the one-dimensional Wasserstein setting and under the stated regularity conditions, the leading regret rate is still governed by the policy-class complexity. Moreover, we provide a minimax lower bound establishing the sharpness of the leading dependence on $N$ and $\mathrm{N\text{-}dim}(Π)$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。