arXiv:2606.00431cs.LG2026-06被引 1

提出更优的贝叶斯强化学习算法,可自适应不同方差场景。

Variance-sensitive Thompson sampling for generalised linear bandits, revisited

  • 基于高斯庞加莱不等式改进传统方法,避免早期分析失效。
  • 证明了方差敏感的后悔上界,性能随方差变化自动调节。
  • 适合高方差环境下的在线决策问题,如推荐系统与广告投放。

我们为随机广义线性马尔可夫带模型中的汤普森采样证明了方差敏感的后悔上界。该分析依赖于一个预热阶段,在此之后通过使用高斯庞加莱不等式控制后悔。这一方法避开了此前基于乐观主义分析在某一点失效的问题。如何在不依赖预热的前提下仍保持相同的方差敏感缩放关系,仍是开放问题,且看似非平凡。

原文摘要 · Abstract (English)

We prove a variance-sensitive regret bound for Thompson sampling in stochastic generalised linear bandits. The argument assumes a warm-up, after which the regret is controlled through using the Gaussian Poincaré inequality. This bypasses the point at which previous optimism-based analyses break down. Removing the warm-up while retaining the same variance-sensitive scaling remains open, and appears nontrivial.

强化学习贝叶斯优化在线决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。