用游戏环境自动生成任务,让智能体学会跨场景空间推理。
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- 用多视角目标统一表示,实现多任务强化学习
- 交互成功率提升4倍,可在未见过环境中零样本泛化
- 适合研究通用视觉运动智能与大规模强化学习的学者
尽管强化学习在语言建模中取得显著进展,但在视觉运动智能体领域尚未完全成功。主要挑战在于模型容易过拟合特定任务或环境,难以获得跨场景的通用行为。本文通过在Minecraft环境中微调强化学习模型,首次证明视觉运动智能体可实现对未见世界的零样本泛化。我们探索了强化学习在提升3D世界中空间推理与交互能力方面的潜力。为解决多任务表示难题,提出跨视角目标设定作为统一的多任务目标空间。为克服人工设计任务的瓶颈,提出在高度可定制的Minecraft环境中自动化生成任务,并构建高效分布式强化学习框架支持大规模训练。实验表明,强化学习使交互成功率提升4倍,并实现跨多样化环境(包括真实世界)的空间推理零样本泛化。研究结果表明,特别是可大规模生成任务的3D仿真环境,对显著提升视觉运动智能体的空间推理能力具有巨大潜力。
原文摘要 · Abstract (English)
While Reinforcement Learning (RL) has achieved remarkable success in language modeling, its triumph hasn't yet fully translated to visuomotor agents. A primary challenge in RL models is their tendency to overfit specific tasks or environments, thereby hindering the acquisition of generalizable behaviors across diverse settings. This paper provides a preliminary answer to this challenge by demonstrating that RL-finetuned visuomotor agents in Minecraft can achieve zero-shot generalization to unseen worlds. Specifically, we explore RL's potential to enhance generalizable spatial reasoning and interaction capabilities in 3D worlds. To address challenges in multi-task RL representation, we analyze and establish cross-view goal specification as a unified multi-task goal space for visuomotor policies. Furthermore, to overcome the significant bottleneck of manual task design, we propose automated task synthesis within the highly customizable Minecraft environment for large-scale multi-task RL training, and we construct an efficient distributed RL framework to support this. Experimental results show RL significantly boosts interaction success rates by $4\times$ and enables zero-shot generalization of spatial reasoning across diverse environments, including real-world settings. Our findings underscore the immense potential of RL training in 3D simulated environments, especially those amenable to large-scale task generation, for significantly advancing visuomotor agents' spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。