arXiv:2603.22760cs.RO2026-03被引 6

让机器人理解空间关系,提升家务操作能力。

SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation

  • 通过多模态输入和辅助任务联合训练增强空间感知。
  • 在13维动作空间下,各类操作成功率显著优于直接模仿学习。
  • 适合研究通用家用机器人与具身智能的开发者参考。

视觉-语言-动作(VLA)模型在机器人控制中展现潜力,但在复杂家庭环境中表现仍不理想。移动操作需对全局场景布局、精细几何结构及高维连续动作进行推理,标准模仿学习难以满足需求。本文提出一种空间接地的VLA建模框架,通过辅助任务联合训练与多模态输入增强来强化感知与表征。方法融合多视角RGB图像、深度信息与短时序历史,同时协同训练多个解码器,从共享的视觉-语言特征中重建全局机器人位置、关节配置、抓取可操作性、目标物相对位姿及分割掩码等中间信号。这些任务提供密集监督,促使主干网络生成具有空间接地性与操作感知能力的隐式表示。在家庭重排任务上的大量评估表明,该方法在抓取、放置、开合等操作上均实现稳定提升,显著优于直接模仿学习。结果表明,通过辅助任务与多模态学习实现空间接地,是推动VLA模型向通用家用机器人演进的重要方向。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained geometry, and high-dimensional continuous actions, making standard imitation learning insufficient. We introduce a framework for learning spatially-grounded VLA models that strengthens perception and representation through auxiliary task co-training and multi-modal input enhancement. Our method addresses the challenge of controlling a 13-dimensional action space involving coordinated base motion, arm articulation, and gripper actuation. To enrich spatial understanding, the model incorporates multi-view RGB observations, depth cues, and short temporal history, providing perspectives of both global scene structure and local manipulation context. To improve representation quality, we co-train auxiliary decoders that reconstruct interpretable intermediate signals - including global robot position, joint configurations, grasp affordances, target-object relative pose, and segmentation masks - from shared visual-language features. These objectives provide dense supervision that encourages the backbone to develop spatially grounded, manipulation-aware latent representations. Through extensive evaluation on home rearrangement tasks, our approach achieves consistent improvements across picking, placing, opening, and closing operations, substantially outperforming direct imitation learning. Our findings suggest that spatial grounding through auxiliary and multi-modal learning provides a strong direction for scaling VLA models toward general-purpose domestic robots.

机器人空间感知多模态具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。