让智能体在开放世界中更高效探索长期回报。
Open-World Reinforcement Learning over Long Short-Term Imagination
- 构建长短期世界模型,扩展想象轨迹长度。
- 在MineDojo上比现有方法提升显著的探索效率。
- 适合研究长期决策与高效探索的学者参考。
在高维开放世界中训练视觉强化学习智能体面临巨大挑战。尽管基于模型的方法通过学习交互式世界模型提升了样本效率,但这些智能体往往存在‘短视’问题,通常仅在短时想象片段上训练。我们认为,开放世界决策的核心挑战在于如何在巨大状态空间中提升探索效率,尤其针对需要考虑长期回报的任务。本文提出LS-Imagine,通过在有限的状态转移步数内扩展想象视野,使智能体能够探索可能带来有利长期反馈的行为。其核心是构建一个‘长短期世界模型’。具体通过模拟目标条件下的跳跃式状态转移,并在单张图像中聚焦特定区域计算对应的能力图(affordance maps),从而将直接的长期价值融入行为学习。该方法在MineDojo基准上显著优于当前最优技术。
原文摘要 · Abstract (English)
Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be "short-sighted", as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。