arXiv:2506.07140stat.MLcs.LG2025-06

在未观测混杂下学习最优分位数策略,理论保证高效稳健。

Quantile-Optimal Policy Learning under Unmeasured Confounding

  • 用工具变量和负控制构建非线性积分方程估计分位数目标
  • 通过极小极大估计与保守策略设计,提升数据覆盖不足下的可靠性
  • 首个在未观测混杂下实现样本高效分位数最优策略的算法

我们研究分位数最优策略学习,目标是找到使回报分布第α分位数最大的策略(α∈(0,1))。聚焦于存在未观测混杂的离线设置,该问题面临三大挑战:(i) 分位数目标对回报分布是非线性的;(ii) 存在未观测混杂;(iii) 离线数据覆盖不足。为此,我们提出一套因果辅助的策略学习方法,在温和条件下具有严格的理论保证。具体地,结合工具变量与负控制,通过求解非线性泛函积分方程来估计分位数目标;采用非参数模型的极小极大估计方法求解这些方程,并构造保守策略以应对覆盖不足问题;最终策略为最大化这些悲观估计值的策略。此外,提出一种新型正则化策略学习方法,更易计算。我们证明,在离线数据覆盖条件较弱的情况下,所学策略达到$ ilde{ ext{O}}(n^{-1/2})$的分位数最优性,其中$ ilde{ ext{O}}(ullet)$省略多项式对数因子。据我们所知,这是首个在存在未观测混杂时实现样本高效的分位数最优策略学习算法。

原文摘要 · Abstract (English)

We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest $α$-quantile for some $α\in (0, 1)$. We focus on the offline setting whose generating process involves unobserved confounders. Such a problem suffers from three main challenges: (i) nonlinearity of the quantile objective as a functional of the reward distribution, (ii) unobserved confounding issue, and (iii) insufficient coverage of the offline dataset. To address these challenges, we propose a suite of causal-assisted policy learning methods that provably enjoy strong theoretical guarantees under mild conditions. In particular, to address (i) and (ii), using causal inference tools such as instrumental variables and negative controls, we propose to estimate the quantile objectives by solving nonlinear functional integral equations. Then we adopt a minimax estimation approach with nonparametric models to solve these integral equations, and propose to construct conservative policy estimates that address (iii). The final policy is the one that maximizes these pessimistic estimates. In addition, we propose a novel regularized policy learning method that is more amenable to computation. Finally, we prove that the policies learned by these methods are $\tilde{\mathscr{O}}(n^{-1/2})$ quantile-optimal under a mild coverage assumption on the offline dataset. Here, $\tilde{\mathscr{O}}(\cdot)$ omits poly-logarithmic factors. To the best of our knowledge, we propose the first sample-efficient policy learning algorithms for estimating the quantile-optimal policy when there exist unmeasured confounding.

策略学习因果推断分位数优化未观测混杂

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。