arXiv:2603.18702cs.LG2026-03

针对有限供应场景优化离线策略学习,提升资源分配效率。

Off-Policy Learning with Limited Supply

  • 基于相对奖励差异选择物品,避免热门项过早耗尽
  • 实验证明在真实与合成数据上优于传统方法
  • 适合库存或预算受限的推荐与广告系统

我们研究上下文老虎机中的离线策略学习(OPL),该问题在推荐系统和在线广告等实际应用中至关重要。传统OPL假设环境无约束,可无限次选择同一物品。但在优惠券发放、电商商品库存等场景中,物品受预算或库存限制,盲目选择当前期望回报最高的物品可能导致其提前耗尽,使未来可能带来更高回报的用户无法使用。这使得在无约束环境下最优的方法在有限供应下表现不佳。我们通过理论分析表明,常规贪婪方法可能无法最大化策略性能,并证明更优策略在有限供应下存在。据此提出新方法OPLS:不只选最高回报物品,而是关注相对于其他用户的相对更高回报项,实现更高效分配。在合成与真实数据集上的实验显示,OPLS在有限供应的上下文老虎机任务中显著优于现有方法。

原文摘要 · Abstract (English)

We study off-policy learning (OPL) in contextual bandits, which plays a key role in a wide range of real-world applications such as recommendation systems and online advertising. Typical OPL in contextual bandits assumes an unconstrained environment where a policy can select the same item infinitely. However, in many practical applications, including coupon allocation and e-commerce, limited supply constrains items through budget limits on distributed coupons or inventory restrictions on products. In these settings, greedily selecting the item with the highest expected reward for the current user may lead to early depletion of that item, making it unavailable for future users who could potentially generate higher expected rewards. As a result, OPL methods that are optimal in unconstrained settings may become suboptimal in limited supply settings. To address the issue, we provide a theoretical analysis showing that conventional greedy OPL approaches may fail to maximize the policy performance, and demonstrate that policies with superior performance must exist in limited supply settings. Based on this insight, we introduce a novel method called Off-Policy learning with Limited Supply (OPLS). Rather than simply selecting the item with the highest expected reward, OPLS focuses on items with relatively higher expected rewards compared to the other users, enabling more efficient allocation of items with limited supply. Our empirical results on both synthetic and real-world datasets show that OPLS outperforms existing OPL methods in contextual bandit problems with limited supply.

离线学习上下文老虎机资源分配推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。