用贝叶斯优化主动选模型,少量在线交互提升离线强化学习性能
Enhancing Offline Model-Based RL via Active Model Selection: A Bayesian Optimization Perspective
- 将模型选择建模为贝叶斯优化问题,引入新型模型诱导核实现高效概率推断
- 仅需1%-2.5%的离线数据量级在线交互,显著提升策略性能
- 适用于数据受限但需高可靠性模型选择的离线强化学习场景
离线模型基强化学习(MBRL)通过预收集数据与学习的动力学模型,可仅凭离线数据训练出高性能策略。为充分发挥其潜力,动力学模型的选择至关重要。然而,传统方法依赖验证或离线评估,因离线强化学习中的分布偏移而精度不足。为此,我们提出BOMS框架,通过贝叶斯优化视角,在仅需少量在线交互的前提下增强离线MBRL的模型选择能力。具体地,将模型选择重构为贝叶斯优化问题,并提出一种理论支持且计算高效的新型模型诱导核。大量实验表明,BOMS在多种强化学习任务上,仅需相当于离线训练数据1%–2.5%的在线交互量,即优于基准方法。
原文摘要 · Abstract (English)
Offline model-based reinforcement learning (MBRL) serves as a competitive framework that can learn well-performing policies solely from pre-collected data with the help of learned dynamics models. To fully unleash the power of offline MBRL, model selection plays a pivotal role in determining the dynamics model utilized for downstream policy learning. However, offline MBRL conventionally relies on validation or off-policy evaluation, which are rather inaccurate due to the inherent distribution shift in offline RL. To tackle this, we propose BOMS, an active model selection framework that enhances model selection in offline MBRL with only a small online interaction budget, through the lens of Bayesian optimization (BO). Specifically, we recast model selection as BO and enable probabilistic inference in BOMS by proposing a novel model-induced kernel, which is theoretically grounded and computationally efficient. Through extensive experiments, we show that BOMS improves over the baseline methods with a small amount of online interaction comparable to only $1\%$-$2.5\%$ of offline training data on various RL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。