让视觉模型具备3D空间感知能力,显著提升机器人任务表现。
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
- 用多视角图像微分渲染增强ViT的3D空间理解能力。
- 在268个任务中超越10种顶尖方法,且训练数据更少。
- 适合关注机器人视觉、多模态学习的研究者使用。
本文提出SPA,一种强调3D空间感知在具身智能中重要性的表示学习框架。该方法通过多视角图像的可微神经渲染,赋予基础视觉变换器(ViT)内在的空间理解能力。我们进行了迄今最全面的具身表示学习评估,涵盖8个仿真环境中的268个任务,覆盖单任务与语言引导多任务场景,采用多种策略。结果表明,SPA持续优于10种以上先进表示方法,包括专为具身智能、视觉任务及多模态应用设计的方法,且所需训练数据更少。此外,我们在真实场景中开展系列实验,验证其实际有效性。最强模型训练耗时超过6000 GPU小时,代码与权重将全部开源,以推动具身表示学习研究。项目页面:https://haoyizhu.github.io/spa/。
原文摘要 · Abstract (English)
In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with intrinsic spatial understanding. We present the most comprehensive evaluation of embodied representation learning to date, covering 268 tasks across 8 simulators with diverse policies in both single-task and language-conditioned multi-task scenarios. The results are compelling: SPA consistently outperforms more than 10 state-of-the-art representation methods, including those specifically designed for embodied AI, vision-centric tasks, and multi-modal applications, while using less training data. Furthermore, we conduct a series of real-world experiments to confirm its effectiveness in practical scenarios. These results highlight the critical role of 3D spatial awareness for embodied representation learning. Our strongest model takes more than 6000 GPU hours to train and we are committed to open-sourcing all code and model weights to foster future research in embodied representation learning. Project Page: https://haoyizhu.github.io/spa/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。