边云协同框架提升工业视觉检测精度与效率
AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection
- 轻量边缘检测与云端多模态大模型协同定位
- 抗遮挡噪声和场景变化,分类准确率显著提升
- 动态调度适配异构设备,兼顾低延迟与高吞吐
多模态大语言模型(MLLM)在语义理解与视觉推理方面表现优异,但在精确目标定位和资源受限的边云部署中仍面临挑战。本文提出AIVD框架,通过轻量级边缘检测器与云端MLLM的协作,实现统一的精准定位与高质量语义生成。为增强云端MLLM对边缘裁剪框噪声和场景变化的鲁棒性,设计了一种高效的微调策略,结合视觉-语义协同增强,显著提升分类准确率与语义一致性。此外,为保障异构边缘设备与动态网络条件下高吞吐与低延迟,提出一种异构资源感知的动态调度算法。实验表明,AIVD显著降低资源消耗,同时提升MLLM分类性能与语义生成质量;所提调度策略在多种场景下均实现更高吞吐与更低延迟。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) demonstrate exceptional capabilities in semantic understanding and visual reasoning, yet they still face challenges in precise object localization and resource-constrained edge-cloud deployment. To address this, this paper proposes the AIVD framework, which achieves unified precise localization and high-quality semantic generation through the collaboration between lightweight edge detectors and cloud-based MLLMs. To enhance the cloud MLLM's robustness against edge cropped-box noise and scenario variations, we design an efficient fine-tuning strategy with visual-semantic collaborative augmentation, significantly improving classification accuracy and semantic consistency. Furthermore, to maintain high throughput and low latency across heterogeneous edge devices and dynamic network conditions, we propose a heterogeneous resource-aware dynamic scheduling algorithm. Experimental results demonstrate that AIVD substantially reduces resource consumption while improving MLLM classification performance and semantic generation quality. The proposed scheduling strategy also achieves higher throughput and lower latency across diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。