让大模型远程决策,边缘设备快速执行,提升机器人导航的实时性与成功率。
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
- 大模型远程做语义判断,边缘小模型高频执行动作,实现异步协同。
- 在6秒延迟下,成功率达92%,比现有方法高40%。
- 适合需要强智能与高响应的边缘机器人应用。
机器人基础模型通过互联网规模的视觉-语言表征实现强大泛化能力,但其巨大的计算开销导致推理延迟过高,成为主要瓶颈:在动态环境中,延迟会破坏控制回路,使强大模型无法安全用于实时部署。我们提出AsyncVLA,一种异步控制框架,将语义推理与反应式执行解耦。受分层控制启发,大型基础模型在远程工作站运行以提供高层指导,而轻量级的机载边缘适配器则以高频持续优化动作。为弥合这两条异步流之间的领域差距,我们引入端到端微调协议和轨迹重加权策略,优先关注动态交互。我们在真实世界的视觉导航任务上评估该方法,通信延迟最高达6秒。AsyncVLA相比最先进基线实现40%更高的成功率,有效连接了大模型的语义智能与边缘机器人所需的反应能力。
原文摘要 · Abstract (English)
Robotic foundation models achieve strong generalization by leveraging internet-scale vision-language representations, but their massive computational cost creates a fundamental bottleneck: high inference latency. In dynamic environments, this latency breaks the control loop, rendering powerful models unsafe for real-time deployment. We propose AsyncVLA, an asynchronous control framework that decouples semantic reasoning from reactive execution. Inspired by hierarchical control, AsyncVLA runs a large foundation model on a remote workstation to provide high-level guidance, while a lightweight, onboard Edge Adapter continuously refines actions at high frequency. To bridge the domain gap between these asynchronous streams, we introduce an end-to-end finetuning protocol and a trajectory re-weighting strategy that prioritizes dynamic interactions. We evaluate our approach on real-world vision-based navigation tasks with communication delays up to 6 seconds. AsyncVLA achieves a 40% higher success rate than state-of-the-art baselines, effectively bridging the gap between the semantic intelligence of large models and the reactivity required for edge robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。