arXiv:2606.20107cs.LG2026-06

提出无需计数的量化均值集成法,实现最小最大最优强化学习。

Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

  • 用量化均值替代传统计数估计不确定性
  • 在有限时域MDP中达到最优方差依赖的遗憾上界
  • 为集成方法探索提供理论支持,适合理论研究者

最优强化学习算法通常依赖精心构造的基于计数的不确定性估计来驱动探索。尽管理论上合理,这类估计在实际场景中难以计算,因而对设计探索启发式帮助有限。与此同时,集成方法虽已作为实用方案出现,却缺乏理论依据。基于最近针对多臂赌博机的集成方法,本文提出一种针对有限时域马尔可夫决策过程(MDPs)的基于分位数的集成方法。该简单、无计数的方法实现了最优的方差依赖遗憾上界,为强化学习中的集成探索提供了理论基础。

原文摘要 · Abstract (English)

Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and therefore offer limited insight for designing exploration heuristics. Meanwhile, ensembling has emerged as a practical approach, but remains without theoretical justification. Building on a recent ensemble-based method for Multi-Armed Bandits, we propose a quantile-based ensemble method for finite-horizon Markov Decision Processes (MDPs). Our simple count-free approach achieves optimal variance-dependent regret bounds, providing theoretical grounding for ensemble-based exploration in RL.

强化学习集成方法理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。