采样任务几乎无需探索,突破传统强化学习认知
Multi-Armed Sampling Problem and the End of Exploration
- 提出多臂采样框架,类比多臂赌博机但聚焦采样
- 理论证明采样只需极少探索,近似最优算法实现低遗憾
- 适用于大模型微调、强化学习等场景,解释为何探索不必要
本文提出多臂采样框架,作为多臂赌博机优化问题的采样对偶。核心动机是严谨分析采样中的探索-利用权衡。我们系统定义了该框架下合理的遗憾度量并建立相应下界,进而提出一种简单算法,实现近似最优遗憾。理论表明,与优化不同,采样几乎无需探索。为连接采样与赌博机,我们引入连续参数族,通过温度参数平滑统一两类问题。我们认为该框架可成为采样研究的基础,如同多臂赌博机在强化学习中的地位。尤其对熵正则化强化学习、预训练模型微调及人类反馈强化学习(RLHF)中的算法收敛性与探索作用提供新理解。
原文摘要 · Abstract (English)
This paper introduces the framework of multi-armed sampling, which serves as the sampling counterpart to the optimization problem of multi-armed bandits. Our primary motivation is to rigorously examine the exploration-exploitation trade-off in the context of sampling. We systematically define plausible notions of regret for this framework and establish corresponding lower bounds. We then propose a simple algorithm that achieves near-optimal regret bounds. Our theoretical results suggest that, in contrast to optimization, sampling barely requires any exploration. To further connect our findings with those of multi-armed bandits, we define a continuous family of problems and associated regret measures that smoothly interpolate and unify multi-armed sampling and multi-armed bandit problems using a temperature parameter. We believe that the multi-armed sampling framework and our findings in this setting can play a foundational role in the study of sampling, including recent neural samplers, much like the role of multi-armed bandits in reinforcement learning. In particular, our work sheds light on the role of exploration (or lack thereof) and the convergence properties of algorithms for entropy-regularized reinforcement learning, fine-tuning of pretrained models and reinforcement learning with human feedback (RLHF).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。