arXiv:2608.02257cs.RO2026-08

用全景视觉提升机器人移动操作的全局空间理解能力

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

论文配图:Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation
图 1 · 摘自论文原文
  • 通过全景编码与融合模块,让模型理解全向视觉信息
  • 在4个真实任务中达到73.4%端到端成功率,比局部视角模型高
  • 适合需要复杂环境交互的移动机械臂系统研究者

移动操作是具身智能的关键能力,使机器人能在开放环境中完成多阶段复杂任务。然而,现有视觉-语言-动作(VLA)策略面临两大挑战:数据层面需协调移动基座与机械臂进行高质量示范采集;模型层面则依赖局部视觉,视野受限影响全局空间理解。为此,我们构建了基于单个VR界面的全身远程操控系统,实现了轮式双臂机器人的协同控制,并收集了5.5小时的真实世界多模态示范数据集。在此基础上,提出PanoVLA——一种全景感知的视觉-语言-动作策略。该模型基于混合变压器架构,引入专用全景编码与融合模块,有效整合全景观测、语言指令与机器人状态以生成动作。在四个真实任务上的评估显示,PanoVLA平均阶段完成率达91.3%,端到端成功率为73.4%,显著优于局部视图基线,证明引入全景空间上下文可提升移动机器人对空间的理解与闭环操作性能。

原文摘要 · Abstract (English)

Mobile manipulation is a key capability for embodied intelligence, enabling robots to accomplish complex multi-stage tasks in open-world environments. However, mobile manipulation poses two key challenges for vision-language-action (VLA) policies: At the data level, the efficient collection of high-quality whole-body demonstrations demands the coordinated control of both the mobile base and the robotic arms; at the model level, existing VLA models predominantly rely on local camera observations, whose limited field of view hinders global spatial understanding. To address these challenges, we develop a whole-body teleoperation system and a panoramic-aware VLA policy. The system enables coordinated control of a wheeled bimanual robot through a single VR interface and supports the acquisition of a real-world mobile manipulation dataset comprising 5.5 hours of multimodal demonstrations. Building upon this dataset, we propose PanoVLA, a panorama-aware vision-language-action policy for mobile bimanual manipulation. Built upon a Mixture-of-Transformers architecture, PanoVLA introduces global spatial context through dedicated panorama encoding and fusion modules, enabling effective integration of panoramic observations with language instructions and robot states for action generation. Evaluation on four real-world mobile manipulation tasks demonstrates that PanoVLA achieves an average stage completion rate of 91.3\% and an end-to-end success rate of 73.4\%, substantially outperforming local-view baselines. These results demonstrate that incorporating panoramic spatial context improves spatial understanding and closed-loop manipulation performance in mobile robots.

移动操作全景感知视觉语言动作远程操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。