解析视觉语言动作模型推理性能,指导实时机器人系统设计
How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
- 构建VLA-Perf分析框架,评估任意模型与系统组合的推理性能
- 发现模型缩放、异步推理等对延迟影响显著,长视频输入成瓶颈
- 揭示部署策略选择:硬件与网络共同决定端到端延迟
视觉语言动作(VLA)模型在具身智能任务中展现出强大能力。然而,将VLA部署于真实机器人需满足严格实时推理要求,而当前对其推理性能的理解仍不充分,主要受限于模型架构与推理系统组合的庞大空间。本文提出核心问题:如何设计未来VLA模型与系统以支持实时推理?为此,我们引入VLA-Perf——一个可分析任意VLA模型与推理系统组合性能的解析模型。通过该工具,我们首次系统研究了VLA推理性能全景。从模型设计角度,考察了模型缩放、架构选择、长视频输入、异步推理及双系统流水线对性能的影响;从部署视角,分析了推理应执行于设备端、边缘服务器还是云端,并揭示硬件能力与网络性能如何共同决定端到端延迟。基于全面评估提炼出15项关键洞见,旨在为未来VLA模型与系统的开发提供实践指导。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently demonstrated impressive capabilities across various embodied AI tasks. While deploying VLA models on real-world robots imposes strict real-time inference constraints, the inference performance landscape of VLA remains poorly understood due to the large combinatorial space of model architectures and inference systems. In this paper, we ask a fundamental research question: How should we design future VLA models and systems to support real-time inference? To address this question, we first introduce VLA-Perf, an analytical performance model that can analyze inference performance for arbitrary combinations of VLA models and inference systems. Using VLA-Perf, we conduct the first systematic study of the VLA inference performance landscape. From a model-design perspective, we examine how inference performance is affected by model scaling, model architectural choices, long-context video inputs, asynchronous inference, and dual-system model pipelines. From the deployment perspective, we analyze where VLA inference should be executed -- on-device, on edge servers, or in the cloud -- and how hardware capability and network performance jointly determine end-to-end latency. By distilling 15 key takeaways from our comprehensive evaluation, we hope this work can provide practical guidance for the design of future VLA models and inference systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。