arXiv:2602.05019cs.LG2026-02

将对抗性上下文下的约束强化学习问题转化为标准强化学习问题

A Simple Reduction Scheme for Constrained Contextual Bandits with Adversarial Contexts via Regression

  • 通过在线回归预测器构建代理奖励,实现约束问题到无约束问题的简化
  • 在对抗性上下文中实现更优的累积违规与遗憾控制,理论分析简洁清晰
  • 适合研究约束强化学习与算法可扩展性的研究人员参考

我们研究对抗性上下文下的约束上下文强化学习(CCB)问题,其中每个动作产生随机奖励并带来随机成本。假设在给定上下文条件下,奖励和成本分别独立来自期望属于已知函数类的分布。在持续设置中,算法在整个时间范围内运行,即使预算耗尽也继续执行。目标是同时控制遗憾与累积约束违规。基于Foster等(2018)提出的SquareCB框架,我们提出一种简单且模块化的算法方案,利用在线回归预言机将约束问题转化为具有自适应代理奖励函数的标准无约束上下文强化学习问题。相比多数关注随机上下文的前期工作,该方法在更通用的对抗性上下文设定下获得更优保证,并具备紧凑透明的分析结构。

原文摘要 · Abstract (English)

We study constrained contextual bandits (CCB) with adversarially chosen contexts, where each action yields a random reward and incurs a random cost. We adopt the standard realizability assumption: conditioned on the observed context, rewards and costs are drawn independently from fixed distributions whose expectations belong to known function classes. We consider the continuing setting, in which the algorithm operates over the entire horizon even after the budget is exhausted. In this setting, the objective is to simultaneously control regret and cumulative constraint violation. Building on the seminal SquareCB framework of Foster et al. (2018), we propose a simple and modular algorithmic scheme that leverages online regression oracles to reduce the constrained problem to a standard unconstrained contextual bandit problem with adaptively defined surrogate reward functions. In contrast to most prior work on CCB, which focuses on stochastic contexts, our reduction yields improved guarantees for the more general adversarial context setting, together with a compact and transparent analysis.

强化学习约束学习在线学习算法设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。