arXiv:2505.00961stat.MLcs.LG2025-05被引 1

提出DOLCE方法,解决上下文博弈中策略评估的偏差问题。

DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects

  • 用滞后上下文构建重要性权重,分离滞后与当前效应
  • 在奖励模型残差条件均值为零时实现偏差抵消
  • 适合高支持不重叠场景,尤其适用于真实数据

上下文博弈中的离策略评估与学习利用记录的交互数据估计并优化目标策略的价值。现有方法大多要求记录策略与目标策略有充分的动作重叠,否则会引入估值与策略梯度估计偏差。为此,我们提出DOLCE(分解离策略评估/学习为滞后与当前效应),仅使用已存储在博弈日志中的滞后上下文,构建滞后边缘化的重要权重,并将目标函数分解为支持鲁棒的滞后校正项和基于模型的当前项。当奖励模型残差在给定滞后上下文和动作下条件均值为零时,可实现偏差抵消。通过多个候选滞后,DOLCE对各滞后估计进行软聚合,并引入基于矩的训练过程,仅用日志中增强的滞后数据即可促进所需不变性。实验表明,DOLCE在离策略评估与学习中均有显著提升,尤其在违反支持比例增加时表现更优。

原文摘要 · Abstract (English)

Off-policy evaluation and learning in contextual bandits use logged interaction data to estimate and optimize the value of a target policy. Most existing methods require sufficient action overlap between the logging and target policies, and violations can bias value and policy gradient estimates. To address this issue, we propose DOLCE (Decomposing Off-policy evaluation/learning into Lagged and Current Effects), which uses only lagged contexts already stored in bandit logs to construct lag-marginalized importance weights and to decompose the objective into a support-robust lagged correction term and a current, model-based term, yielding bias cancellation when the reward-model residual is conditionally mean-zero given the lagged context and action. With multiple candidate lags, DOLCE softly aggregates lag-specific estimates, and we introduce a moment-based training procedure that promotes the desired invariance using only logged lag-augmented data. We show that DOLCE is unbiased in an idealized setting and yields consistent and asymptotically normal estimates with cross-fitting under standard conditions. Our experiments demonstrate that DOLCE achieves substantial improvements in both off-policy evaluation and learning, particularly as the proportion of individuals who violate support increases.

离策略评估上下文博弈偏差修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。