arXiv:2605.06190cs.LG2026-05

在对抗性上下文中实现带预算约束的强化学习,兼顾收益与成本控制。

Constrained Contextual Bandits with Adversarial Contexts

论文配图:Constrained Contextual Bandits with Adversarial Contexts
图 1 · 摘自论文原文
  • 通过在线回归预言机将约束问题转化为代理奖励的无约束问题。
  • 在对抗性上下文场景下,实现更优的累积损失与预算违规控制。
  • 算法简洁透明,适用于持续运行的现实场景,如广告投放。

我们研究了具有对抗性上下文的预算约束上下文老虎机问题,其中每个动作产生随机奖励并消耗随机成本。采用标准可实现性假设:在给定观测上下文的条件下,奖励和成本分别独立地从其期望属于已知函数类的固定分布中抽取。聚焦于持续运行场景,即即使累计成本预算耗尽,算法仍继续运行。目标是同时控制遗憾值和预算违规程度。基于Foster等人[2018]的开创性SquareCB框架,我们提出一个简单且模块化的框架,利用在线回归预言机将约束问题转化为带有自适应定义代理奖励函数的标准无约束上下文老虎机问题。与先前仅关注随机上下文的工作不同,我们的方法在更一般的对抗性上下文设定下获得了更优的保证,并提供了一种高效、紧凑且清晰的分析方式。

原文摘要 · Abstract (English)

We study budget-constrained contextual bandits with adversarial contexts, where each action yields a random reward and incurs a random cost. We adopt the standard realizability assumption: conditioned on the observed context, rewards and costs are drawn independently from fixed distributions whose expectations belong to known function classes. We focus on the continuing setting, in which the algorithm operates over the entire horizon even after the budget for cumulative cost is exhausted. In this setting, the objective is to simultaneously control regret and the violation of the budget constraint. Building on the seminal $\mathsf{SquareCB}$ framework of Foster et al. [2018], we propose a simple and modular framework that leverages online regression oracles to reduce the constrained problem to a standard unconstrained contextual bandit problem with adaptively defined surrogate reward functions. In contrast to prior works, which focus on stochastic contexts, our reduction yields improved guarantees for more general adversarial contexts, together with an efficient algorithm with a compact and transparent analysis.

强化学习上下文老虎机对抗性环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。