提出新方法提升强化学习中探索与利用的平衡效率。
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
- 用变分贝叶斯与序列蒙特卡洛结合实现高效规划。
- 单块显卡下可扩展更大规划预算,样本与运行效率更优。
- 适合追求高数据效率的强化学习研究者使用。
在强化学习中,如何最优地权衡探索与利用是实现任务求解最大数据效率的关键。贝叶斯最优智能体虽能达成此目标,但其信念状态推断与规划通常不可行。尽管深度学习有助于缓解计算复杂性,现有方法仍需高昂训练成本。本文提出一种变分框架,用于贝叶斯自适应马尔可夫决策过程中的学习与规划,融合变分信念学习、序列蒙特卡洛规划与元强化学习。在单块GPU环境下,所提方法VariBASeD展现出对更大规划预算的良好可扩展性,在样本效率与运行效率上均优于先前方法。
原文摘要 · Abstract (English)
Optimally trading-off exploration and exploitation is the holy grail of reinforcement learning as it promises maximal data-efficiency for solving any task. Bayes-optimal agents achieve this, but obtaining the belief-state and performing planning are both typically intractable. Although deep learning methods can greatly help in scaling this computation, existing methods are still costly to train. To accelerate this, this paper proposes a variational framework for learning and planning in Bayes-adaptive Markov decision processes that coalesces variational belief learning, sequential Monte-Carlo planning, and meta-reinforcement learning. In a single-GPU setup, our new method VariBASeD exhibits favorable scaling to larger planning budgets, improving sample- and runtime-efficiency over prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。