arXiv:2511.20066cs.LG2025-11NeurIPS被引 7

SOMBRL通过不确定性乐观原则,实现高效探索的模型化强化学习。

SOMBRL: Scalable and Optimistic Model-Based RL

  • 基于不确定性乐观策略,动态调整探索与奖励权重。
  • 在非线性系统中证明了亚线性遗憾,适用于多种任务场景。
  • 适合需要高效探索的复杂控制任务,尤其硬件部署表现优异。

我们解决模型化强化学习(MBRL)中高效探索的挑战,其中系统动态未知,智能体需直接从在线交互中学习。提出可扩展且乐观的模型化强化学习(SOMBRL),基于面对不确定性时的乐观原则。SOMBRL学习一个感知不确定性的动力学模型,并贪婪地最大化外在奖励与认知不确定性加权和。SOMBRL兼容任意策略优化器或规划器,在常见正则性假设下,证明其在非线性动态系统中具有亚线性遗憾,覆盖有限时域、折扣无限时域及非周期设置。此外,SOMBRL提供了灵活可扩展的原理性探索方案。我们在状态空间与视觉控制环境中评估SOMBRL,其在所有任务中均优于基线方法。还在动态遥控车硬件上验证,SOMBRL超越现有最先进方法,体现了原理性探索在MBRL中的优势。

原文摘要 · Abstract (English)

We address the challenge of efficient exploration in model-based reinforcement learning (MBRL), where the system dynamics are unknown and the RL agent must learn directly from online interactions. We propose Scalable and Optimistic MBRL (SOMBRL), an approach based on the principle of optimism in the face of uncertainty. SOMBRL learns an uncertainty-aware dynamics model and greedily maximizes a weighted sum of the extrinsic reward and the agent's epistemic uncertainty. SOMBRL is compatible with any policy optimizers or planners, and under common regularity assumptions on the system, we show that SOMBRL has sublinear regret for nonlinear dynamics in the (i) finite-horizon, (ii) discounted infinite-horizon, and (iii) non-episodic settings. Additionally, SOMBRL offers a flexible and scalable solution for principled exploration. We evaluate SOMBRL on state-based and visual-control environments, where it displays strong performance across all tasks and baselines. We also evaluate SOMBRL on a dynamic RC car hardware and show SOMBRL outperforms the state-of-the-art, illustrating the benefits of principled exploration for MBRL.

强化学习模型化探索策略硬件部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。