arXiv:2508.12638cs.CVcs.AI2025-08被引 5

用延迟的云端大模型输出做边缘实时推理的上下文,提升响应速度与准确率。

edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer

  • 将云端大模型的延迟输出作为历史上下文,指导边缘小模型实时推理
  • 在四个数据集上实测,显著降低延迟并保持高精度
  • 适合对实时性与准确性要求高的自动驾驶、人机交互场景

视觉语言模型(VLMs)正广泛应用于自动驾驶和人机交互等实时场景,需基于精准感知快速可靠地响应。现有系统多采用云边协同架构,如分割式大视觉语言模型(LVLMs)或大小模型间任务卸载策略,但这些方法无法应对云延迟波动,且未充分利用延迟但准确的LVLM输出。本文提出一种新型云边协同范式——上下文传递(Context Transfer),将LVLM的延迟输出作为历史上下文,为边缘小模型(SVLM)推理提供实时指导。基于此,我们设计了edgeVLM,引入上下文替换与视觉聚焦模块,优化历史文本输入并增强视觉定位一致性。在四个数据集上的三个实时视觉语言推理任务中,实验验证了该框架的有效性。该范式为未来更高效、低延迟感知的VLM系统提供了新思路。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly deployed in real-time applications such as autonomous driving and human-computer interaction, which demand fast and reliable responses based on accurate perception. To meet these requirements, existing systems commonly employ cloud-edge collaborative architectures, such as partitioned Large Vision-Language Models (LVLMs) or task offloading strategies between Large and Small Vision-Language Models (SVLMs). However, these methods fail to accommodate cloud latency fluctuations and overlook the full potential of delayed but accurate LVLM responses. In this work, we propose a novel cloud-edge collaborative paradigm for VLMs, termed Context Transfer, which treats the delayed outputs of LVLMs as historical context to provide real-time guidance for SVLMs inference. Based on this paradigm, we design edgeVLM, which incorporates both context replacement and visual focus modules to refine historical textual input and enhance visual grounding consistency. Extensive experiments on three real-time vision-lanuage reasoning tasks across four datasets demonstrate the effectiveness of the proposed framework. The new paradigm lays the groundwork for more effective and latency-aware collaboration strategies in future VLM systems.

云边协同视觉语言模型实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。