让机器人动作直接在摄像头视角中预测,提升真实场景泛化能力
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

- 将动作预测从机械臂基座坐标系转为摄像头坐标系,统一多视角输出
- 在模拟与真实任务中均显著提升成功率和跨视角泛化性能
- 无需修改原有模型结构,可直接接入现有视觉语言动作系统
视觉-语言-动作(VLA)模型在真实环境中的泛化能力常受观测空间与动作空间不一致的影响。尽管训练数据来自多种摄像头视角,但模型通常在机械臂基座坐标系中预测末端执行器位姿,导致空间错配。为此,本文提出观察中心型视觉-语言-动作(OC-VLA)框架,将动作预测直接锚定在摄像头观测空间。利用摄像头外参矩阵,将末端执行器位姿从机械臂基座坐标系转换至摄像头坐标系,实现异构视角下的预测目标统一。该轻量级、即插即用策略有效提升了感知与动作的对齐性,显著增强模型对摄像头视角变化的鲁棒性。方法可兼容现有VLA架构,无需重大改动。在模拟与真实机器人操作任务上的全面评估表明,OC-VLA加速收敛,提高任务成功率,并改善跨视角泛化能力。代码将公开。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces. Although training data are collected from diverse camera perspectives, the models typically predict end-effector poses within the robot base coordinate frame, resulting in spatial inconsistencies. To mitigate this limitation, we introduce the Observation-Centric VLA (OC-VLA) framework, which grounds action predictions directly in the camera observation space. Leveraging the camera's extrinsic calibration matrix, OC-VLA transforms end-effector poses from the robot base coordinate system into the camera coordinate system, thereby unifying prediction targets across heterogeneous viewpoints. This lightweight, plug-and-play strategy ensures robust alignment between perception and action, substantially improving model resilience to camera viewpoint variations. The proposed approach is readily compatible with existing VLA architectures, requiring no substantial modifications. Comprehensive evaluations on both simulated and real-world robotic manipulation tasks demonstrate that OC-VLA accelerates convergence, enhances task success rates, and improves cross-view generalization. The code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。