用多模态大模型提升边缘设备目标检测的准确性与效率
Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- 通过指令微调让多模态大模型生成场景描述,实现语义增强
- 在低光和遮挡场景中,延迟降低79%,计算成本减少70%仍保持精度
- 边缘-云协同框架根据置信度自动选择是否调用云端语义指导
传统目标检测方法在低光照和严重遮挡等复杂场景下因缺乏高层语义理解而性能下降。为此,本文提出一种基于自适应引导的语义增强边缘-云协同目标检测方法,利用多模态大语言模型(MLLM)实现准确率与效率的平衡。具体而言,首先通过指令微调使MLLM生成结构化场景描述;随后设计自适应映射机制,将语义信息动态转化为边缘检测器的参数调整信号,实现实时语义增强。在边缘-云协同推理框架中,系统根据置信度自动决定是否调用云端语义引导或直接输出边缘检测结果。实验表明,该方法在复杂场景中显著提升检测性能:在低光与高度遮挡场景下,延迟降低超过79%,计算成本减少70%,同时维持高精度。
原文摘要 · Abstract (English)
Traditional object detection methods face performance degradation challenges in complex scenarios such as low-light conditions and heavy occlusions due to a lack of high-level semantic understanding. To address this, this paper proposes an adaptive guidance-based semantic enhancement edge-cloud collaborative object detection method leveraging Multimodal Large Language Models (MLLM), achieving an effective balance between accuracy and efficiency. Specifically, the method first employs instruction fine-tuning to enable the MLLM to generate structured scene descriptions. It then designs an adaptive mapping mechanism that dynamically converts semantic information into parameter adjustment signals for edge detectors, achieving real-time semantic enhancement. Within an edge-cloud collaborative inference framework, the system automatically selects between invoking cloud-based semantic guidance or directly outputting edge detection results based on confidence scores. Experiments demonstrate that the proposed method effectively enhances detection accuracy and efficiency in complex scenes. Specifically, it can reduce latency by over 79% and computational cost by 70% in low-light and highly occluded scenes while maintaining accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。