arXiv:2409.18286cs.CVcs.AI2024-09综述被引 31

评测多模态大模型在交通目标检测中的表现与挑战

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

  • 系统梳理多模态大模型在交通场景下的应用与技术脉络
  • 实测三类真实交通任务,揭示模型在安全属性提取等方面的性能表现
  • 提出未来研究方向,适合关注智能交通AI落地的研究者参考

本研究旨在全面回顾并实证评估多模态大语言模型(MLLMs)和大视觉模型(VLMs)在交通系统目标检测中的应用。首先,阐述了MLLMs在交通应用中的潜在优势,并综述现有研究中MLLMs的技术进展,重点分析其在不同交通场景下目标检测的有效性与局限性。其次,构建了交通场景端到端目标检测的分类体系,并展望未来发展方向。在此基础上,针对道路安全属性提取、关键安全事件检测及热成像图像视觉推理三个真实交通问题,开展实证分析,评估MLLMs性能,揭示其优势与改进空间。最后,讨论了当前实际应用中面临的限制与挑战,为该领域未来研究与开发提供路线图。

原文摘要 · Abstract (English)

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first fold, we provide a background about the potential benefits of MLLMs in transportation applications and conduct a comprehensive review of current MLLM technologies in previous studies. We highlight their effectiveness and limitations in object detection within various transportation scenarios. The second fold involves providing an overview of the taxonomy of end-to-end object detection in transportation applications and future directions. Building on this, we proposed empirical analysis for testing MLLMs on three real-world transportation problems that include object detection tasks namely, road safety attributes extraction, safety-critical event detection, and visual reasoning of thermal images. Our findings provide a detailed assessment of MLLM performance, uncovering both strengths and areas for improvement. Finally, we discuss practical limitations and challenges of MLLMs in enhancing object detection in transportation, thereby offering a roadmap for future research and development in this critical area.

多模态模型目标检测智能交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。