arXiv:2601.14921cs.ROcs.AI2026-01被引 3

将视觉语言模型部署在边缘计算节点,实现机器人实时感知。

Vision-Language Models on the Edge for Real-Time Robotic Perception

  • 用WebRTC流传输多模态数据至边缘节点执行推理
  • 边缘部署使端到端延迟降低5%,接近云端精度
  • 小模型可在资源受限下实现亚秒级响应,适合实时场景

视觉语言模型(VLMs)支持机器人多模态感知与交互,但其在真实系统中的部署受限于延迟、本地算力不足以及云端卸载带来的隐私风险。6G时代的开放无线接入网(Open RAN)与多接入边缘计算(MEC)为解决这些问题提供了可能,通过将计算贴近数据源实现高效处理。本文以Unitree G1人形机器人作为实体测试平台,研究在ORAN/MEC架构上部署VLM的可行性。设计基于WebRTC的多模态数据流传输管道,对比评估了在边缘与云端部署的LLaMA-3.2-11B-Vision-Instruct模型在实时条件下的表现。结果表明,边缘部署在保持近似云端精度的同时,将端到端延迟降低5%。进一步测试了专为资源受限环境优化的Qwen2-VL-2B-Instruct小模型,其响应时间低于1秒,延迟降幅超一半,但牺牲一定准确性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) enable multimodal reasoning for robotic perception and interaction, but their deployment in real-world systems remains constrained by latency, limited onboard resources, and privacy risks of cloud offloading. Edge intelligence within 6G, particularly Open RAN and Multi-access Edge Computing (MEC), offers a pathway to address these challenges by bringing computation closer to the data source. This work investigates the deployment of VLMs on ORAN/MEC infrastructure using the Unitree G1 humanoid robot as an embodied testbed. We design a WebRTC-based pipeline that streams multimodal data to an edge node and evaluate LLaMA-3.2-11B-Vision-Instruct deployed at the edge versus in the cloud under real-time conditions. Our results show that edge deployment preserves near-cloud accuracy while reducing end-to-end latency by 5\%. We further evaluate Qwen2-VL-2B-Instruct, a compact model optimized for resource-constrained environments, which achieves sub-second responsiveness, cutting latency by more than half but at the cost of accuracy.

边缘计算视觉语言模型机器人感知实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。