arXiv:2603.02271cs.PFcs.AI2026-03被引 2

发现边缘AI中动作生成占75%延迟,制约视觉语言模型落地

Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures

  • 分析主流边缘硬件上视觉语言动作模型的执行瓶颈
  • 动作生成阶段耗时占端到端延迟75%,是主要性能瓶颈
  • 适合关注边缘AI部署与硬件优化的研究者和工程师

视觉-语言-动作(VLA)模型是机器人和具身智能在边缘计算中的关键新兴工作负载。随着模型规模扩大,其能力显著提升,但必须本地部署以满足实时应用的严苛延迟要求。本文在Nvidia Jetson Orin和Thor两代边缘硬件上评估了VLA性能。基于当前最先进的MolmoAct-7B模型,我们识别出主要执行瓶颈:高达75%的端到端延迟由内存受限的动作生成阶段占用。通过解析建模与仿真,我们预测了将模型扩展至100B参数规模所需的硬件条件。同时探讨了高带宽内存技术及存内计算(PIM)作为未来边缘系统中具身智能的潜在优化路径。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are an emerging class of workloads critical for robotics and embodied AI at the edge. As these models scale, they demonstrate significant capability gains, yet they must be deployed locally to meet the strict latency requirements of real-time applications. This paper characterizes VLA performance on two generations of edge hardware, viz. the Nvidia Jetson Orin and Thor platforms. Using MolmoAct-7B, a state-of-the-art VLA model, we identify a primary execution bottleneck: up to 75% of end-to-end latency is consumed by the memory-bound action-generation phase. Through analytical modeling and simulations, we project the hardware requirements for scaling to 100B parameter models. We also explore the impact of high-bandwidth memory technologies and processing-in-memory (PIM) as promising future pathways in edge systems for embodied AI.

边缘计算VLA模型性能瓶颈硬件优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。