arXiv:2412.06139cs.LGcs.SY2024-12中稿 · as a poster presen…被引 2

用世界模型不确定性约束探索,提升强化学习效率

Bounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning Algorithm

  • 结合软探索与内在动机,基于世界模型不确定性控制探索范围
  • 在8组实验中6组取得最高分,加速模型收敛
  • 适合奖励函数严格、需高效探索的现实场景

深度强化学习在真实世界应用中的瓶颈之一是如何高效探索环境并收集有意义的状态转移。本文提出一种名为有界探索的新方法,融合了‘软’探索与内在动机探索。该方法显著提升了Soft Actor-Critic算法及其基于模型扩展版本的性能,加快了收敛速度。在8个实验中,有6个实验取得了最高得分。当原始奖励函数具有严格意义时,有界探索为引入内在动机提供了替代方案。

原文摘要 · Abstract (English)

One of the bottlenecks preventing Deep Reinforcement Learning algorithms (DRL) from real-world applications is how to explore the environment and collect informative transitions efficiently. The present paper describes bounded exploration, a novel exploration method that integrates both 'soft' and intrinsic motivation exploration. Bounded exploration notably improved the Soft Actor-Critic algorithm's performance and its model-based extension's converging speed. It achieved the highest score in 6 out of 8 experiments. Bounded exploration presents an alternative method to introduce intrinsic motivations to exploration when the original reward function has strict meanings.

强化学习探索策略世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。