arXiv:2601.22830cs.CVcs.RO2026-01

评测10个大模型在恶劣条件下的2D目标检测能力,发现语义推理强但定位不准。

A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions

  • 用PeSOTIF数据集系统评估十种大视觉语言模型的检测性能
  • 顶级模型召回率比YOLOv5高25%以上,但在几何精度上仍落后于专用检测器
  • 适合用于自动驾驶安全验证,可补充传统检测器的不足

可靠环境感知仍是自动驾驶安全运行的主要障碍。功能安全性(SOTIF)关注因感知不足带来的风险,尤其在恶劣条件下传统检测器常失效。尽管大视觉语言模型(LVLM)展现出出色的语义推理能力,其在安全关键的2D目标检测中的量化效果尚未充分探索。本文基于专为长尾交通场景和环境退化设计的PeSOTIF数据集,对十种代表性LVLM进行了系统评估,并与两种专用检测器——基于锚点的YOLOv5和基于变换器的RT-DETRv4进行对比。实验结果揭示关键权衡:顶级LVLM(如Gemini 3)在自然视觉退化下召回率较YOLOv5提升超25%,并接近RT-DETRv4表现;而专用检测器在手工构造扰动下仍保持几何精度优势。这些发现凸显了语义推理与几何回归的互补性,支持将LVLM作为面向SOTIF的自动驾驶系统中的高层安全验证工具。

原文摘要 · Abstract (English)

Reliable environmental perception remains one of the main obstacles for safe operation of automated vehicles. Safety of the Intended Functionality (SOTIF) concerns safety risks from perception insufficiencies, particularly under adverse conditions where conventional detectors often falter. While Large Vision-Language Models (LVLMs) demonstrate promising semantic reasoning, their quantitative effectiveness for safety-critical 2D object detection is underexplored. This paper presents a systematic evaluation of ten representative LVLMs using the PeSOTIF dataset, a benchmark specifically curated for long-tail traffic scenarios and environmental degradations. Performance is quantitatively compared against two specialized detectors: the anchor-based YOLOv5 and the transformer-based RT-DETRv4. Experimental results reveal a critical trade-off: top-performing LVLMs (e.g., Gemini 3) surpass the YOLOv5 in recall by over 25% and closely match RT-DETRv4 under natural visual degradation, while specialized detectors retain an advantage in geometric precision for handcrafted perturbations. These findings highlight the complementary strengths of semantic reasoning versus geometric regression, supporting the use of LVLMs as high-level safety validators in SOTIF-oriented automated driving systems.

目标检测自动驾驶大模型SOTIF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。