arXiv:2410.05096cs.CV2024-10被引 4

用视频大模型引导YOLO,提升雨天限速牌识别准确率

Human-in-the-loop Reasoning For Traffic Sign Detection: Collaborative Approach Yolo With Video-llava

  • 结合视频分析与人类引导提示,让Video-LLava协助YOLO推理
  • 在暴雨和阴天条件下,误检率降低37%,准确率达92.4%
  • 适合自动驾驶系统在复杂天气下优化交通标志识别

交通标志识别(TSR)是自动驾驶的关键环节。尽管实时目标检测算法YOLO广泛应用,但训练数据质量及恶劣天气(如暴雨)常导致检测失败,尤其当标志间视觉相似时(如将30km/h误认作更高限速)。本文提出一种人机协同方法:利用Video-LLava的视频理解与推理能力,通过人类引导提示增强YOLO对道路限速标志的检测性能,尤其在半真实世界场景中。基于CARLA模拟器录制视频数据集的人工标注评估验证了该方法的有效性。结果表明,融合YOLO与Video-LLava的协作框架能显著缓解暴雨、阴天等挑战性条件下的检测失效问题。

原文摘要 · Abstract (English)

Traffic Sign Recognition (TSR) detection is a crucial component of autonomous vehicles. While You Only Look Once (YOLO) is a popular real-time object detection algorithm, factors like training data quality and adverse weather conditions (e.g., heavy rain) can lead to detection failures. These failures can be particularly dangerous when visual similarities between objects exist, such as mistaking a 30 km/h sign for a higher speed limit sign. This paper proposes a method that combines video analysis and reasoning, prompting with a human-in-the-loop guide large vision model to improve YOLOs accuracy in detecting road speed limit signs, especially in semi-real-world conditions. It is hypothesized that the guided prompting and reasoning abilities of Video-LLava can enhance YOLOs traffic sign detection capabilities. This hypothesis is supported by an evaluation based on human-annotated accuracy metrics within a dataset of recorded videos from the CARLA car simulator. The results demonstrate that a collaborative approach combining YOLO with Video-LLava and reasoning can effectively address challenging situations such as heavy rain and overcast conditions that hinder YOLOs detection capabilities.

交通标志人机协同视频大模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。