arXiv:2509.19096cs.CVcs.SE2025-09中稿 · presentation at th…被引 3

用多模态大模型零样本检测交通事故,提升智能监控效率。

Investigating Traffic Accident Detection Using Multimodal Large Language Models

  • 结合YOLO、Deep SORT、SAM构建增强提示,提升模型理解力。
  • Pixtral在无训练下表现最佳,F1达0.71,召回率83%。
  • 适合智能交通、自动驾驶领域研究者参考。

交通安全隐患仍是全球性难题,及时准确的事故检测对降低风险和快速响应至关重要。基于基础设施的视觉传感器可实现连续实时监测,直接从摄像头图像中自动识别事故。本研究探索多模态大语言模型(MLLMs)在零样本条件下,利用基础设施摄像头图像检测并描述交通肇事事件的能力,减少对大量标注数据的依赖。主要贡献包括:(1) 使用CARLA仿真环境生成的DeepAccident数据集评估MLLMs,解决真实、多样化基础设施事故数据稀缺问题;(2) 对比Gemini 1.5与2.0、Gemma 3与Pixtral模型在未微调情况下的事故识别与描述能力;(3) 在提示中集成YOLO目标检测、Deep SORT多目标跟踪与Segment Anything(SAM)实例分割技术,提升模型准确性和可解释性。关键结果表明,Pixtral表现最优,F1为0.71,召回率达83%;而Gemini系列在增强提示后精度提升(如Gemini 1.5达90%),但F1与召回显著下降;Gemma 3则表现最均衡,指标波动最小。结果表明,将MLLMs与先进视觉分析技术融合具有巨大潜力,可显著提升其在真实世界自动化交通监控中的应用价值。

原文摘要 · Abstract (English)

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for continuous real-time monitoring, facilitating automated detection of accidents directly from captured images. This research investigates the zero-shot capabilities of multimodal large language models (MLLMs) for detecting and describing traffic accidents using images from infrastructure cameras, thus minimizing reliance on extensive labeled datasets. Main contributions include: (1) Evaluation of MLLMs using the simulated DeepAccident dataset from CARLA, explicitly addressing the scarcity of diverse, realistic, infrastructure-based accident data through controlled simulations; (2) Comparative performance analysis between Gemini 1.5 and 2.0, Gemma 3 and Pixtral models in accident identification and descriptive capabilities without prior fine-tuning; and (3) Integration of advanced visual analytics, specifically YOLO for object detection, Deep SORT for multi-object tracking, and Segment Anything (SAM) for instance segmentation, into enhanced prompts to improve model accuracy and explainability. Key numerical results show Pixtral as the top performer with an F1-score of 0.71 and 83% recall, while Gemini models gained precision with enhanced prompts (e.g., Gemini 1.5 rose to 90%) but suffered notable F1 and recall losses. Gemma 3 offered the most balanced performance with minimal metric fluctuation. These findings demonstrate the substantial potential of integrating MLLMs with advanced visual analytics techniques, enhancing their applicability in real-world automated traffic monitoring systems.

交通监控多模态模型零样本检测视觉分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。