arXiv:2603.26360cs.RO2026-03被引 4

让视觉语言机器人系统在真实场景中实现高速、平滑、精准运行

Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate

  • 整合校准、规划控制与学习方法,优化端到端执行速度
  • 实测机器人操作速度接近人类日常水平,逼近轻量机械臂硬件极限
  • 适合关注机器人实时推理部署的开发者与研究者

在将视觉语言模型(VLA)应用于实际机器人任务时,执行速度至关重要。此前工作arXiv:2510.26742分析了如何在GPU上加速神经计算,但未解决实际机器人部署问题。本报告描述了一套实用技术,实现了在真实世界任务中以极高速度运行VLA驱动机器人的端到端效果,兼具准确性与灵巧性。技术栈涵盖校准、规划与控制,以及基于学习的方法以确定最优执行速度。实验显示,机器人在任务中的执行速度已达到与日常人类操作相当的水平,并逼近我们所用轻量级机械臂的硬件极限。未加速视频与推理轨迹详见https://dexmal.github.io/realtime-vla-v2/

原文摘要 · Abstract (English)

In deployment of the VLA models to real-world robotic tasks, execution speed matters. In previous work arXiv:2510.26742 we analyze how to make neural computation of VLAs on GPU fast. However, we leave the question of how to actually deploy the VLA system on the real robots open. In this report we describe a set of practical techniques to achieve the end-to-end result of running a VLA-driven robot at an impressive speed in real world tasks that require both accuracy and dexterity. The stack of technology ranges across calibration, planning & control, and learning based method to identify optimal execution speed. In the tasks we show, the robot even executes in a speed on par with casual human operation and approaching the hardware limit of our lightweight arm. The unaccelerated videos and inference traces are provided in https://dexmal.github.io/realtime-vla-v2/.

机器人实时推理视觉语言模型部署优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。