动态环境中应对对抗性约束的强化学习新算法
Bandits in Flux: Adversarial Constraints in Dynamic Environments
- 基于对偶优化与梯度估计改进在线镜像下降
- 实现次线性动态遗憾与次线性约束违反
- 适合需要实时适应变化环境的决策系统
我们研究了在时间变化约束下运行的对抗性多臂老虎机问题,该场景源于众多现实应用。为应对这一复杂设定,提出一种新颖的原-对偶算法,通过引入合适的梯度估计器和有效的约束处理机制扩展了在线镜像下降方法。理论分析证明所提策略具有次线性动态遗憾和次线性约束违反。实验表明,该算法在遗憾和约束违反两个指标上均达到当前最优性能。
原文摘要 · Abstract (English)
We investigate the challenging problem of adversarial multi-armed bandits operating under time-varying constraints, a scenario motivated by numerous real-world applications. To address this complex setting, we propose a novel primal-dual algorithm that extends online mirror descent through the incorporation of suitable gradient estimators and effective constraint handling. We provide theoretical guarantees establishing sublinear dynamic regret and sublinear constraint violation for our proposed policy. Our algorithm achieves state-of-the-art performance in terms of both regret and constraint violation. Empirical evaluations demonstrate the superiority of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。