arXiv:2607.11498cs.ROcs.AI2026-07被引 1

用机器人坐标系的点图解决视觉语言动作模型的视角不匹配问题

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

论文配图:See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建机器人坐标系下的3D点图,让视觉输入与动作空间对齐
  • 在RoboCasa上使成功率提升,尤其在训练未见视角下优势更明显
  • 适配现有2D视觉模型,无需修改架构,适合真实机器人部署

视觉-语言-动作(VLA)模型从视觉观测和语言指令中预测机器人动作。这些动作定义在机器人的3D坐标系中,但大多数VLA在相机坐标系下观察场景,造成观测与动作定义帧之间的不匹配。固定视角下该问题可被策略通过记忆单一映射解决,但当大规模数据集包含多种相机设置时,策略需跨视角泛化,问题加剧。本文提出机器人中心点图(robot-centric pointmaps),其像素存储场景点在机器人坐标系中的3D坐标。点图在保留预训练2D VLA所需的密集H×W网格结构的同时,提供机器人坐标系下的3D几何信息,可无缝集成到现有VLA中,仅需最小架构改动。在RoboCasa数据集上,点图提升了pi0.5和SmolVLA性能,并优于代表性相机视角和3D感知基线。真实机器人实验显示,当摄像头移至训练中未见位置时,点图策略相比纯RGB策略的优势进一步扩大。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training.

视觉动作点图机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。