通过裁剪小倾向得分提升离线策略学习的稳定性与性能。
Offline Policy Learning with Weight Clipping and Heaviside Composite Optimization
- 用裁剪阈值控制小倾向得分,降低策略价值估计方差。
- 设计分段优化框架,解决裁剪带来的非连续性问题。
- 适合有历史数据、需稳定策略优化的个性化决策场景。
离线策略学习旨在利用历史数据学习最优个性化决策规则。在标准的先估计后优化框架中,基于重加权的方法(如逆倾向评分加权或双重稳健估计器)被广泛用于生成无偏的策略价值估计。然而,当某些处理的倾向得分较小时,这些方法在策略价值估计上会面临高方差问题,可能误导下游策略优化,导致学习到的策略表现不佳。本文系统地提出一种基于权重裁剪估计器的离线策略学习算法,该估计器通过选择最小化策略价值估计均方误差(MSE)的裁剪阈值来截断小倾向得分。针对线性策略,我们通过将问题重构为Heaviside复合优化问题,解决了由权重裁剪引发的双层与不连续目标问题,提供了严格的计算框架。随后采用渐进整数规划方法高效求解该优化问题,使实际策略学习成为可能。我们建立了所提算法次优性的上界,揭示了通过权重裁剪降低策略价值估计的MSE,如何提升策略学习性能。
原文摘要 · Abstract (English)
Offline policy learning aims to use historical data to learn an optimal personalized decision rule. In the standard estimate-then-optimize framework, reweighting-based methods (e.g., inverse propensity weighting or doubly robust estimators) are widely used to produce unbiased estimates of policy values. However, when the propensity scores of some treatments are small, these reweighting-based methods suffer from high variance in policy value estimation, which may mislead the downstream policy optimization and yield a learned policy with inferior value. In this paper, we systematically develop an offline policy learning algorithm based on a weight-clipping estimator that truncates small propensity scores via a clipping threshold chosen to minimize the mean squared error (MSE) in policy value estimation. Focusing on linear policies, we address the bilevel and discontinuous objective induced by weight-clipping-based policy optimization by reformulating the problem as a Heaviside composite optimization problem, which provides a rigorous computational framework. The reformulated policy optimization problem is then solved efficiently using the progressive integer programming method, making practical policy learning tractable. We establish an upper bound for the suboptimality of the proposed algorithm, which reveals how the reduction in MSE of policy value estimation, enabled by our proposed weight-clipping estimator, leads to improved policy learning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。