Dreamer 4在模拟世界中离线训练,无需真实交互即可在Minecraft中获得钻石。
Training Agents Inside of Scalable World Models
- 用高效Transformer和快捷强制目标实现实时交互推理
- 仅用少量数据学习通用动作条件,从大量无标签视频中提取知识
- 首次纯离线完成超2万步操作的Minecraft钻石获取任务
世界模型从视频中学习通用知识,并在想象中模拟经验以训练行为,为智能体发展提供新路径。然而,以往世界模型难以准确预测复杂环境中的物体交互。我们提出Dreamer 4,一种可在快速准确的世界模型内通过强化学习训练的可扩展智能体。在复杂游戏Minecraft中,该世界模型精准预测物体交互与游戏机制,性能远超先前模型。通过快捷强制目标与高效Transformer架构,模型实现单张GPU上的实时交互推理。此外,模型仅需少量数据即可学习通用动作条件,从而从多样化未标注视频中提取大部分知识。我们提出了仅基于离线数据获取Minecraft钻石的挑战,契合机器人等领域中因安全与效率限制无法进行真实交互的实际需求。该任务需从原始像素中选择超过20,000次鼠标键盘操作序列。通过在想象中学习行为,Dreamer 4是首个完全依赖离线数据在Minecraft中成功获取钻石的智能体。本工作提供了可扩展的想象训练范式,推动了智能体的发展。
原文摘要 · Abstract (English)
World models learn general knowledge from videos and simulate experience for training behaviors in imagination, offering a path towards intelligent agents. However, previous world models have been unable to accurately predict object interactions in complex environments. We introduce Dreamer 4, a scalable agent that learns to solve control tasks by reinforcement learning inside of a fast and accurate world model. In the complex video game Minecraft, the world model accurately predicts object interactions and game mechanics, outperforming previous world models by a large margin. The world model achieves real-time interactive inference on a single GPU through a shortcut forcing objective and an efficient transformer architecture. Moreover, the world model learns general action conditioning from only a small amount of data, allowing it to extract the majority of its knowledge from diverse unlabeled videos. We propose the challenge of obtaining diamonds in Minecraft from only offline data, aligning with practical applications such as robotics where learning from environment interaction can be unsafe and slow. This task requires choosing sequences of over 20,000 mouse and keyboard actions from raw pixels. By learning behaviors in imagination, Dreamer 4 is the first agent to obtain diamonds in Minecraft purely from offline data, without environment interaction. Our work provides a scalable recipe for imagination training, marking a step towards intelligent agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。