对比边缘与云端GPU部署视觉语言动作模型的性能表现
Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
- 测试5种VLA模型在不同硬件上的表现,分析架构设计影响
- 边缘设备在低功耗下仍可达到甚至超越旧数据中心GPU性能
- 高吞吐方案可在不损失精度前提下实现,适合资源受限场景
视觉-语言-动作(VLA)模型已成为机器人控制的强大通用策略,但其在不同模型架构和硬件平台上的性能扩展规律及功耗预算仍不明确。本文评估了五种代表性VLA模型——涵盖先进基线和两种新提出的架构——在边缘与数据中心GPU平台上的表现。基于LIBERO基准测试,测量准确率以及延迟、吞吐量和峰值内存使用等系统级指标,覆盖不同边缘功耗约束与高性能数据中心GPU配置。结果揭示出显著的缩放趋势:(1) 架构选择如动作标记化方式与模型主干尺寸显著影响吞吐量与内存占用;(2) 功耗受限的边缘设备呈现非线性性能退化,部分配置表现可匹配甚至超过旧款数据中心GPU;(3) 高吞吐变体可在不显著牺牲准确率的前提下实现。这些发现为在多样化部署约束下选择与优化VLA提供了可操作洞见。本研究挑战了当前认为数据中心硬件在机器人推理中始终占优的假设。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as powerful generalist policies for robotic control, yet their performance scaling across model architectures and hardware platforms, as well as their associated power budgets, remain poorly understood. This work presents an evaluation of five representative VLA models -- spanning state-of-the-art baselines and two newly proposed architectures -- targeting edge and datacenter GPU platforms. Using the LIBERO benchmark, we measure accuracy alongside system-level metrics, including latency, throughput, and peak memory usage, under varying edge power constraints and high-performance datacenter GPU configurations. Our results identify distinct scaling trends: (1) architectural choices, such as action tokenization and model backbone size, strongly influence throughput and memory footprint; (2) power-constrained edge devices exhibit non-linear performance degradation, with some configurations matching or exceeding older datacenter GPUs; and (3) high-throughput variants can be achieved without significant accuracy loss. These findings provide actionable insights when selecting and optimizing VLAs across a range of deployment constraints. Our work challenges current assumptions about the superiority of datacenter hardware for robotic inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。