用2D图像直接预测机器人末端位姿,训练更高效且适配多种机器人。
cVLA: Towards Efficient Camera-Space VLAs
- 基于视觉语言模型,直接从2D图像推断机械臂末端坐标。
- 轻量设计下仍能生成可执行的轨迹点,实测具备强仿真到现实迁移能力。
- 支持深度图与演示条件控制,适合快速部署于真实机器人系统。
视觉-语言-动作(VLA)模型为复杂机器人操作任务提供了有力框架,但训练成本高昂。本文提出一种新型VLA方法,利用视觉语言模型在2D图像上的优异表现,直接从图像帧坐标中推断机器人末端执行器位姿。不同于以往输出低层控制信号的VLA模型,本模型预测轨迹航点,兼具训练效率高和机器人形态无关性。尽管结构轻量,其采用的下一步词预测架构仍能有效学习有意义且可执行的机器人轨迹。我们进一步探索了深度图像的潜力、推理时解码策略及演示条件动作生成等未被充分利用的技术。模型在模拟数据上训练,并展现出强大的仿真到现实迁移能力。通过结合模拟与真实数据评估,验证了其在真实机器人系统上的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive performance of Vision Language Models (VLMs) on 2D images to directly infer robot end-effector poses in image frame coordinates. Unlike prior VLA models that output low-level controls, our model predicts trajectory waypoints, making it both more efficient to train and robot embodiment agnostic. Despite its lightweight design, our next-token prediction architecture effectively learns meaningful and executable robot trajectories. We further explore the underutilized potential of incorporating depth images, inference-time techniques such as decoding strategies, and demonstration-conditioned action generation. Our model is trained on a simulated dataset and exhibits strong sim-to-real transfer capabilities. We evaluate our approach using a combination of simulated and real data, demonstrating its effectiveness on a real robotic system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。