用变化驱动内在动机提升世界模型智能体在稀疏奖励环境中的表现
World Model Agents with Change-Based Intrinsic Motivation
- 基于环境变化设计内在激励,增强世界模型的探索能力
- 在复杂环境Crafter中提升收益,在简单环境Minigrid中反而降低表现
- 适合研究稀疏奖励下探索机制与任务目标对齐问题的研究者
稀疏奖励环境给强化学习带来重大挑战,因反馈稀缺。内在动机与迁移学习成为应对策略。本文将基于变化的探索迁移(CBET)方法适配至世界模型算法如DreamerV3,对比其在Crafter和Minigrid两个稀疏奖励环境下的表现。实验表明,无预训练时,CBET可提升DreamerV3在Crafter中的回报,但在Minigrid中导致性能下降;迁移学习实验显示,以内在奖励预训练的DreamerV3未能快速适应外在奖励目标。结果表明,CBET在复杂环境中有正向作用,但在简单环境中可能因探索行为与任务目标不一致而损害性能。
原文摘要 · Abstract (English)
Sparse reward environments pose a significant challenge for reinforcement learning due to the scarcity of feedback. Intrinsic motivation and transfer learning have emerged as promising strategies to address this issue. Change Based Exploration Transfer (CBET), a technique that combines these two approaches for model-free algorithms, has shown potential in addressing sparse feedback but its effectiveness with modern algorithms remains understudied. This paper provides an adaptation of CBET for world model algorithms like DreamerV3 and compares the performance of DreamerV3 and IMPALA agents, both with and without CBET, in the sparse reward environments of Crafter and Minigrid. Our tabula rasa results highlight the possibility of CBET improving DreamerV3's returns in Crafter but the algorithm attains a suboptimal policy in Minigrid with CBET further reducing returns. In the same vein, our transfer learning experiments show that pre-training DreamerV3 with intrinsic rewards does not immediately lead to a policy that maximizes extrinsic rewards in Minigrid. Overall, our results suggest that CBET provides a positive impact on DreamerV3 in more complex environments like Crafter but may be detrimental in environments like Minigrid. In the latter case, the behaviours promoted by CBET in DreamerV3 may not align with the task objectives of the environment, leading to reduced returns and suboptimal policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。