arXiv:2508.19294cs.CVcs.AI2025-08综述被引 51

大模型让目标检测更懂语义,提升定位与泛化能力。

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

  • 融合视觉与语言的多模态模型增强上下文理解
  • 在多种场景中实现精准定位与分割,表现超越传统方法
  • 适合研究目标检测与机器人感知的开发者参考

视觉语言大模型(LVLM)通过融合自然语言处理与计算机视觉技术,显著提升了目标检测的适应性、上下文推理能力和泛化性能,超越传统架构。本文系统梳理了当前LVLM在目标检测中的进展,分三步展开:首先解析其工作原理,阐述如何结合NLP与CV实现检测与定位革新;其次介绍近期模型的架构创新、训练范式与输出灵活性,强调其在复杂上下文理解上的优势;再分析视觉与文本信息融合策略,展示其推动检测与定位方法演进的成果。文中包含多组可视化案例,验证模型在不同场景下的有效性,并对比其在实时性、适应性与复杂度方面优于传统深度学习系统。基于此,预计未来LVLM将在目标检测中达到甚至超过现有方法水平。同时指出当前模型存在若干局限,提出改进建议并绘制未来发展路线图。结论认为,LVLM的进展将持续深刻影响目标检测与机器人应用。

原文摘要 · Abstract (English)

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This in-depth review presents a structured exploration of the state-of-the-art in LVLMs, systematically organized through a three-step research review process. First, we discuss the functioning of vision language models (VLMs) for object detection, describing how these models harness natural language processing (NLP) and computer vision (CV) techniques to revolutionize object detection and localization. We then explain the architectural innovations, training paradigms, and output flexibility of recent LVLMs for object detection, highlighting how they achieve advanced contextual understanding for object detection. The review thoroughly examines the approaches used in integration of visual and textual information, demonstrating the progress made in object detection using VLMs that facilitate more sophisticated object detection and localization strategies. This review presents comprehensive visualizations demonstrating LVLMs' effectiveness in diverse scenarios including localization and segmentation, and then compares their real-time performance, adaptability, and complexity to traditional deep learning systems. Based on the review, its is expected that LVLMs will soon meet or surpass the performance of conventional methods in object detection. The review also identifies a few major limitations of the current LVLM modes, proposes solutions to address those challenges, and presents a clear roadmap for the future advancement in this field. We conclude, based on this study, that the recent advancement in LVLMs have made and will continue to make a transformative impact on object detection and robotic applications in the future.

目标检测多模态大模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。