提出一种无需调参的批量集成方法,显著提升随机多臂赌博机的在线学习性能。
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
- 采用批量集成策略,通过单一参数控制探索与利用平衡。
- 理论证明在随机多臂赌博机上实现近最优后悔率。
- 适用于对参数敏感度低、追求稳定表现的在线学习场景。
高效权衡探索与利用是在线强化学习中的核心挑战。现有方法通常通过精确估计模型不确定性并遵循所谓的乐观模型来实现。受实际集成方法启发,本文提出一种简单且新颖的批量集成方案,理论上可为随机多臂赌博机(MAB)达到近最优后悔率。关键在于该算法仅有一个参数——批次数量,其取值不依赖于损失分布的尺度或方差等特性。我们通过合成基准测试验证了该算法的有效性,补充了理论结果。
原文摘要 · Abstract (English)
Efficiently trading off exploration and exploitation is one of the key challenges in online Reinforcement Learning (RL). Most works achieve this by carefully estimating the model uncertainty and following the so-called optimistic model. Inspired by practical ensemble methods, in this work we propose a simple and novel batch ensemble scheme that provably achieves near-optimal regret for stochastic Multi-Armed Bandits (MAB). Crucially, our algorithm has just a single parameter, namely the number of batches, and its value does not depend on distributional properties such as the scale and variance of the losses. We complement our theoretical results by demonstrating the effectiveness of our algorithm on synthetic benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。