用正交图像生成提升机器人指令理解的泛化能力
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
- 将多视角观测转为正交视图,实现输入视角不变性
- 在Arnold和Colosseum上实现超40%泛化性能提升
- 适合需要跨场景泛化的机器人视觉语言动作研究
我们提出OG-VLA,一种结合视觉语言动作模型(VLAs)泛化能力与3D感知策略鲁棒性的新架构。针对自然语言指令与一个或多个RGBD观测映射到准静态机器人动作的挑战,该方法将来自不同视角的输入观测反投影为点云,并渲染成规范正交视图,确保输入输出空间的一致性与视角无关性。这些正交视图通过视觉主干、大语言模型(LLM)和图像扩散模型处理,生成编码末端执行器位置与朝向的图像。在Arnold和Colosseum基准测试中,相较现有方法实现超过40%的相对性能提升,同时保持对已知环境的稳健表现。实际部署仅需3至5次示范即可适应新任务。
原文摘要 · Abstract (English)
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural language instructions and one or more RGBD observations to quasi-static robot actions. 3D-aware robot policies achieve state-of-the-art performance on precise robot manipulation tasks, but struggle with generalization to unseen instructions, scenes, and objects. On the other hand, VLAs excel at generalizing across instructions and scenes, but can be sensitive to camera and robot pose variations. We leverage prior knowledge embedded in language and vision foundation models to improve generalization of 3D-aware keyframe policies. OG-VLA unprojects input observations from diverse views into a point cloud which is then rendered from canonical orthographic views, ensuring input view invariance and consistency between input and output spaces. These canonical views are processed with a vision backbone, a Large Language Model (LLM), and an image diffusion model to generate images that encode the next position and orientation of the end-effector on the input scene. Evaluations on the Arnold and Colosseum benchmarks demonstrate state-of-the-art generalization to unseen environments, with over 40% relative improvements while maintaining robust performance in seen settings. We also show real-world adaption in 3 to 5 demonstrations along with strong generalization. Videos and resources at https://og-vla.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。