arXiv:2410.04080cs.LG2024-10被引 1

提出新分析方法,让跨上下文强化学习算法在未知分布下仍能保证高概率近优误差。

High Probability Bound for Cross-Learning Contextual Bandits with Unknown Context Distributions

  • 利用不同周期间的弱依赖结构,改进原算法的理论分析。
  • 首次证明该算法在未知上下文分布时仍具备高概率近优后悔界。
  • 适合关注理论严谨性与实际鲁棒性的强化学习研究者。

针对在线竞价和睡眠老虎机等场景,研究跨上下文情境下的强化学习问题:学习者可观察到所有可能上下文中对应动作的损失,而不仅限于当前回合的上下文。考虑损失由对抗方选择,上下文独立同分布采样。此前工作假设上下文分布已知,而本研究关注更现实的未知分布情形。我们对Schneider和Zimmert(2023)提出的算法进行深入分析,揭示其原始分析仅能导出期望后悔界,但通过引入周期间弱依赖结构并改进鞅不等式,首次证明该算法可在高概率意义下实现近优后悔性能。关键在于克服了传统鞅工具的适用限制。

原文摘要 · Abstract (English)

Motivated by applications in online bidding and sleeping bandits, we examine the problem of contextual bandits with cross learning, where the learner observes the loss associated with the action across all possible contexts, not just the current round's context. Our focus is on a setting where losses are chosen adversarially, and contexts are sampled i.i.d. from a specific distribution. This problem was first studied by Balseiro et al. (2019), who proposed an algorithm that achieves near-optimal regret under the assumption that the context distribution is known in advance. However, this assumption is often unrealistic. To address this issue, Schneider and Zimmert (2023) recently proposed a new algorithm that achieves nearly optimal expected regret. It is well-known that expected regret can be significantly weaker than high-probability bounds. In this paper, we present a novel, in-depth analysis of their algorithm and demonstrate that it actually achieves near-optimal regret with high probability. There are steps in the original analysis by Schneider and Zimmert (2023) that lead only to an expected bound by nature. In our analysis, we introduce several new insights. Specifically, we make extensive use of the weak dependency structure between different epochs, which was overlooked in previous analyses. Additionally, standard martingale inequalities are not directly applicable, so we refine martingale inequalities to complete our analysis.

强化学习上下文带宽高概率边界博弈优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。