arXiv:2502.10826cs.LGcs.IT2025-02被引 3

提出新方法提升离线上下文多臂老虎机的策略选择与学习效果

Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing

  • 用投注机制构建新置信区间,自适应方差优化策略选择
  • 小样本下冻结策略显著降低方差,性能优于现有方法
  • 适用于数据有限场景,尤其适合低样本率下的稳定学习

我们研究离线上下文多臂老虎机中的策略选择与学习问题,目标是利用固定行为策略收集的数据,训练出最大化回报的策略。本文贡献有二:其一,提出一种基于投注机制的新置信界方法,应用于逆倾向权重序列,实现比以往工作更优、自适应方差的理论保证;其二,提出一种通用优化目标条件,平衡偏差与方差,其中一种特殊形式称为'冻结',在小样本场景下能有效降低方差。分析表明该方法可达到当前最优理论保障。实验结果显示,所提选择方法优于现有方法,且冻结策略在小样本条件下表现更佳。

原文摘要 · Abstract (English)

We consider off-policy selection and learning in contextual bandits, where the learner aims to select or train a reward-maximizing policy using data collected by a fixed behavior policy. Our contribution is two-fold. First, we propose a novel off-policy selection method that leverages a new betting-based confidence bound applied to an inverse propensity weight sequence. Our theoretical analysis reveals that this method achieves a significantly improved, variance-adaptive guarantee over prior work. Second, we propose a novel and generic condition on the optimization objective for off-policy learning that strikes a different balance between bias and variance. One special case, which we call freezing, tends to induce low variance, which is preferred in small-data regimes. Our analysis shows that it matches the best existing guarantees. In our empirical study, our selection method outperforms existing methods, and freezing exhibits improved performance in small-sample regimes.

上下文多臂离线学习策略选择小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。