arXiv:2503.06313cs.CVcs.AI2025-03被引 11

融合深度学习与多模态大模型,提升自动驾驶交通标志识别与车道检测可靠性。

Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection

  • 用ResNet-50、YOLOv8等模型实现高精度交通标志识别。
  • 在夜间和雨天等恶劣条件下仍保持93%以上车道检测准确率。
  • 无需预训练,轻量级框架可推理复杂路况,适合实际部署。

自动驾驶车辆需在复杂动态环境中可靠识别交通标志并稳定检测车道以确保安全行驶。本文提出融合先进深度学习与多模态大语言模型(MLLM)的综合道路感知方法。针对交通标志识别,系统评估了ResNet-50、YOLOv8和RT-DETR,分别取得99.8%、98.0%和96.6%的准确率;对于车道检测,采用基于CNN的分割方法结合多项式拟合,在良好条件下表现优异。此外,提出一种轻量级、基于指令微调的多模态大模型框架,仅使用小而多样数据集即可完成训练,无需初始预训练。该框架能有效处理多种车道类型、复杂交叉口及汇流区,在恶劣条件下通过推理显著提升检测可靠性:帧级整体准确率(FRM)达53.87%,问答整体准确率(QNS)达82.83%;晴天车道检测准确率达99.6%,夜间为93.0%;雨天因视线遮挡导致车道不可见时推理准确率为88.4%,道路退化场景下为95.6%。所提框架大幅增强自动驾驶感知可靠性,适用于多样且挑战性强的道路场景。

原文摘要 · Abstract (English)

Autonomous vehicles (AVs) require reliable traffic sign recognition and robust lane detection capabilities to ensure safe navigation in complex and dynamic environments. This paper introduces an integrated approach combining advanced deep learning techniques and Multimodal Large Language Models (MLLMs) for comprehensive road perception. For traffic sign recognition, we systematically evaluate ResNet-50, YOLOv8, and RT-DETR, achieving state-of-the-art performance of 99.8% with ResNet-50, 98.0% accuracy with YOLOv8, and achieved 96.6% accuracy in RT-DETR despite its higher computational complexity. For lane detection, we propose a CNN-based segmentation method enhanced by polynomial curve fitting, which delivers high accuracy under favorable conditions. Furthermore, we introduce a lightweight, Multimodal, LLM-based framework that directly undergoes instruction tuning using small yet diverse datasets, eliminating the need for initial pretraining. This framework effectively handles various lane types, complex intersections, and merging zones, significantly enhancing lane detection reliability by reasoning under adverse conditions. Despite constraints in available training resources, our multimodal approach demonstrates advanced reasoning capabilities, achieving a Frame Overall Accuracy (FRM) of 53.87%, a Question Overall Accuracy (QNS) of 82.83%, lane detection accuracies of 99.6% in clear conditions and 93.0% at night, and robust performance in reasoning about lane invisibility due to rain (88.4%) or road degradation (95.6%). The proposed comprehensive framework markedly enhances AV perception reliability, thus contributing significantly to safer autonomous driving across diverse and challenging road scenarios.

自动驾驶多模态大模型车道检测交通标志识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。