让机器人视觉语言模型在低功耗下实时运行,自动分配计算任务
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints

- 设计统一架构抽象,动态分配模型计算到设备与边缘端
- 实测能耗降低236%,在20赫兹频率下仍满足实时性要求
- 适合对能效和延迟敏感的机器人部署场景
视觉-语言-动作(VLA)模型是具身智能的潜在基础,但其高推理成本给机器人系统部署带来挑战。设备端受限于算力和能耗预算,难以同时满足实时控制与能效需求;而将推理任务卸载至边缘服务器又受网络波动影响,存在不可预测的延迟风险。为此,本文提出EcoVLA,一种面向VLA模型的自适应设备-边缘协同推理框架,旨在满足实时约束下的系统级能效最大化。EcoVLA首先构建跨不同VLA范式的统一阶段级抽象,建立与架构无关的协同推理设计空间;进而建立联合设备-边缘-网络的延迟与能耗预测模型,实现候选方案的毫秒级运行时评估。在此基础上,持续选择满足实时约束的能效最优方案,并适应网络与系统状态的动态变化。此外,引入轻量级中间张量传输机制,降低跨设备协作的通信开销。实验结果表明,在20赫兹动作输出频率约束下,相较于现有协同推理方法,EcoVLA可提升系统能效达236%,并在动态网络与边缘负载条件下始终满足服务等级协议(SLO)要求。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute capacity and energy budgets, struggling to simultaneously satisfy real-time control and energy efficiency requirements. Alternatively, offloading the inference workload to an edge server is susceptible to fluctuations in system conditions, introducing unpredictable latency risks. Device-edge co-inference offers a promising solution, but systematic research tailored to VLA models remains scarce, particularly a unified co-inference framework that jointly addresses real-time constraints and system-level energy efficiency. Thus, we propose EcoVLA, an adaptive device-edge co-inference framework for VLA models that maximizes system energy efficiency under real-time constraints. EcoVLA first introduces a unified stage-level abstraction over different VLA paradigms, establishing an architecture-agnostic co-inference design space. It then formulates a joint device-edge-network latency and energy prediction model to enable rapid runtime evaluation of candidate co-inference schemes. Building on this, EcoVLA continuously selects the energy-optimal scheme satisfying real-time constraints with millisecond-level overhead, adapting to runtime variations in network and system states. Furthermore, EcoVLA incorporates a lightweight transmission mechanism for inter-stage intermediate tensors to reduce the communication overhead incurred by cross-device collaboration. Experimental results across VLA models show that EcoVLA improves system energy efficiency by up to 236% over existing co-inference approaches under a 20 Hz action output frequency constraint, while consistently maintaining SLO satisfaction under dynamic network and edge workload conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。