针对智能体视觉语言动作模型,提出高效边缘云协同推理框架
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
- 设计冗余感知与兼容性优化的边缘云分割策略
- 实测推理速度提升1.73倍,延迟仅增加5%~7%
- 适合对实时性要求高的机器人视觉任务
视觉语言动作(VLA)模型是具身智能主流,但推理成本高。边缘-云协同(ECC)推理通过减轻边缘设备负担以满足实时需求,但现有方法对VLA模型效果不佳:一是环境导向的分割易受视觉噪声干扰;二是忽略具身任务特有的逐步冗余,破坏动作连续性。为此,本文提出新型ECC推理框架RAPID,针对该框架开发了具体实现。实验表明,该方法在仅引入5%~7%额外开销的前提下,实现最高1.73倍的推理加速。
原文摘要 · Abstract (English)
Vision Language Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) inference offers an effective fix by easing edge-device computing pressure to meet real-time needs. However, existing ECC frameworks are suboptimal for VLA models due to two challenges: (1) Mainstream environment-oriented edge-cloud partitioning methods are susceptible to interference from visual noise; (2) Existing edge-cloud partitioning methods overlook the step-wise redundancy unique to embodied tasks, thereby disrupting the physical continuity of motion. To address these issues, we propose a novel ECC inference framework, termed RAPID. Specifically, we developed an implementation tailored to the proposed framework. Experiments demonstrate this achieves a speedup of up to 1.73x with only 5%~7% overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。