让云端模型适应延迟,实现移动机器人快速响应
Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization

- 云端用延迟观测训练慢变任务特征,边缘融合最新特征与本地视觉
- 40步延迟下成功率仍达63.8%~78.0%,远超基线(最高6.4%)
- 适合需要高响应速度的机器人系统,尤其支持大模型云端部署
在移动机器人上部署千亿参数的视觉-语言-动作(VLA)策略引发系统矛盾:语义推理依赖云端GPU,而闭环控制需本地响应,受网络延迟和抖动影响。现有分层异步策略虽提升吞吐量,但其慢路径表示仍可能过时,或需显式调度与延迟提示。我们提出CloudEdgeVLA,将时间错位视为表示学习问题。云端VLA将延迟观测编码为缓慢变化的任务特征,边缘轻量头结合最新云端特征与当前本地视觉。训练中,当前帧与随机延迟帧均配以相同动作目标,在新旧路径间学习。该目标促使云端表示保留任务级信息,边缘路径提供状态敏感修正。在四个LIBERO基准上,当存在40步均匀延迟时,CloudEdgeVLA成功率达63.8%~78.0%,而VLASH最高仅6.4%,单路径基线最高3.0%。通过移除控制回路中的阻塞同步,该设计为可扩展的VLA部署提供了实用路径,使云端模型可不断增长,而边缘计算保持轻量且响应迅速。
原文摘要 · Abstract (English)
Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Existing hierarchical and asynchronous policies improve throughput, but their slow-path representations can still arrive stale or require explicit scheduling and delay cues. We introduce CloudEdgeVLA, a cloud-edge policy that treats temporal misalignment as a representation-learning problem. A cloud VLA encodes delayed observations into slowly varying task features, while a lightweight edge head combines the latest available cloud feature with current local vision. During training, current and randomly delayed frames are paired with the same current action target in fresh and stale paths. This objective encourages the cloud representation to preserve task-level information while the edge path supplies state-sensitive corrections. Across four LIBERO suites, CloudEdgeVLA retains 63.8--78.0% success with a 40-step uniform-delay window, whereas VLASH reaches at most 6.4% and the evaluated single-path baselines at most 3.0%. By removing blocking synchronization from the control loop, the design offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。