对比主流视觉语言模型在交通工程任务中的表现,发现其分类效果好但定位仍有不足。
Evaluation and Comparison of Visual Language Models for Transportation Engineering Problems
- 用零样本提示测试多个开源与闭源视觉语言模型
- 分类任务表现接近传统CNN,检测定位精度待提升
- 为交通场景智能分析提供可复用的评估基准
近期视觉语言模型(VLM)在图像理解任务中展现出巨大潜力。本研究探索了最先进的VLM模型在基于视觉的交通工程任务中的应用,包括图像分类和目标检测。图像分类任务涵盖拥堵检测与裂缝识别,目标检测任务则聚焦头盔违规行为识别。我们采用CLIP、BLIP、OWL-ViT、Llava-Next等开源模型以及GPT-4o等闭源模型,通过零样本提示(zero-shot prompting)进行评估。该方法无需特定任务的标注数据或微调,即可完成任务。结果显示,在图像分类任务中,VLM模型性能与基准卷积神经网络(CNN)模型相当;但在目标定位任务上仍需改进。本研究系统评估了当前主流VLM模型的优劣,为未来优化及大规模应用提供了基准参考。
原文摘要 · Abstract (English)
Recent developments in vision language models (VLM) have shown great potential for diverse applications related to image understanding. In this study, we have explored state-of-the-art VLM models for vision-based transportation engineering tasks such as image classification and object detection. The image classification task involves congestion detection and crack identification, whereas, for object detection, helmet violations were identified. We have applied open-source models such as CLIP, BLIP, OWL-ViT, Llava-Next, and closed-source GPT-4o to evaluate the performance of these state-of-the-art VLM models to harness the capabilities of language understanding for vision-based transportation tasks. These tasks were performed by applying zero-shot prompting to the VLM models, as zero-shot prompting involves performing tasks without any training on those tasks. It eliminates the need for annotated datasets or fine-tuning for specific tasks. Though these models gave comparative results with benchmark Convolutional Neural Networks (CNN) models in the image classification tasks, for object localization tasks, it still needs improvement. Therefore, this study provides a comprehensive evaluation of the state-of-the-art VLM models highlighting the advantages and limitations of the models, which can be taken as the baseline for future improvement and wide-scale implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。