arXiv:2606.30456cs.ROcs.CV2026-06

将视觉-语言-动作模型迁移到真实机械臂,发现系统级控制比模型改进更关键。

Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform

论文配图:Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform
图 1 · 摘自论文原文
  • 构建了从数据采集到部署的完整实机流程,兼容RLDS标准
  • 实测发现模型在仿真中表现好但在真实机器人上行为不稳定
  • 成功部署依赖于动作语义、坐标系等系统细节的精准对齐

本项目探究了近期视觉-语言-动作(VLA)模型能否从受控研究基准可靠地迁移至真实世界机器人平台——具体为UR5e机械臂。工作涵盖真实机器人数据采集、符合RLDS格式的数据集工程、OpenVLA及OpenVLA-OFT模型的微调与部署,并对动作表示和控制接口进行了系统验证。成果包括:(i) 完整的实机数据采集流水线,(ii) 符合RLDS标准的数据集转换流程,(iii) 初步的VLA模型微调与推理基础设施,(iv) 基于真实机器人实验的结构化观察。这些共同建立了超越仿真的学习型操作系统的可复现评估框架。实证结果显示,离线指标良好与闭环行为不稳定之间存在持续差距;该差距不能仅归因于模型局限,更受动作语义、坐标系约定、模态间时间对齐、图像预处理一致性及数据集覆盖度与质量的强烈影响。由此得出核心结论:真实场景中VLA系统的成功部署,取决于对整个数据-模型-控制链路的精确控制,而非模型容量的增量提升。项目将基于VLA的机器人技术从以模型为中心的问题,重构为系统级挑战,揭示了实机任务执行的困难,并提供了清晰、实验验证的可靠部署条件。

原文摘要 · Abstract (English)

This project investigates whether recent Vision-Language-Action (VLA) models can be transferred from controlled research benchmarks to a real-world robotic platform, specifically a UR5e manipulator, in a reproducible and operationally meaningful manner. The work integrates real-robot data acquisition, dataset engineering (compatible with the RLDS format), and the fine-tuning and deployment of OpenVLA and OpenVLA-OFT models, with systematic validation of action representations and control interfaces. The project resulted in several foundational assets: (i) a complete real-robot data acquisition pipeline, (ii) a dataset conversion workflow aligned with RLDS standards, (iii) an initial fine-tuning and inference infrastructure for VLA models, and (iv) a structured set of experimental observations grounded in real-robot trials. These elements collectively establish a reproducible framework for evaluating learning-based manipulation systems beyond simulation. Empirically, the experiments reveal a consistent gap between promising offline indicators and unstable closed-loop behavior on the physical system: this gap cannot be attributed solely to model limitations, it is strongly influenced by action semantics, coordinate frame conventions, temporal alignment between modalities, image preprocessing consistency, and dataset coverage and quality. These observations lead to a key interpretation: the successful deployment of VLA systems in real-world settings depends less on incremental improvements in model capacity and more on precise control of the entire data-model-control pipeline. The project reframes VLA-based robotics from a primarily model-centric challenge to a system-level problem; it highlights the difficulty of running robust task execution on the real robot and provides a clear, experimentally grounded understanding of the conditions required for reliable deployment.

机器人视觉语言动作实机部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。