arXiv:2501.05566cs.CVcs.AI2025-01被引 40

用CLIP模型实现自动驾驶动态场景理解,实时精准识别复杂路况。

Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding

  • 基于CLIP模型构建实时动态场景检索系统,支持边缘设备部署。
  • 在本田场景数据集上达91.1%最高F1分数,优于零样本GPT-4o。
  • 适合研究自动驾驶视觉理解、ADAS系统优化的开发者与工程师。

场景理解对提升驾驶安全、生成以人为本的自动驾驶决策解释,以及利用人工智能进行行车视频回溯分析至关重要。本研究基于对比语言-图像预训练(CLIP)模型,开发了一种可优化为边缘设备实时部署的动态场景检索系统。在包含约80小时标注驾驶视频的Honda Scenes Dataset上进行帧级分析,结果表明CLIP模型在自然语言监督下具备强大的视觉概念学习能力。微调ViT-L/14和ViT-B/32模型显著提升了场景分类性能,最高F1得分为91.1%。该系统展现出快速精准的场景识别能力,满足高级驾驶辅助系统(ADAS)的关键需求。研究表明,CLIP模型能为动态场景理解与分类提供可扩展、高效的框架,推动更智能、更安全、更具上下文感知能力的自动驾驶技术发展。

原文摘要 · Abstract (English)

Scene understanding is essential for enhancing driver safety, generating human-centric explanations for Automated Vehicle (AV) decisions, and leveraging Artificial Intelligence (AI) for retrospective driving video analysis. This study developed a dynamic scene retrieval system using Contrastive Language-Image Pretraining (CLIP) models, which can be optimized for real-time deployment on edge devices. The proposed system outperforms state-of-the-art in-context learning methods, including the zero-shot capabilities of GPT-4o, particularly in complex scenarios. By conducting frame-level analysis on the Honda Scenes Dataset, which contains a collection of about 80 hours of annotated driving videos capturing diverse real-world road and weather conditions, our study highlights the robustness of CLIP models in learning visual concepts from natural language supervision. Results also showed that fine-tuning the CLIP models, such as ViT-L/14 and ViT-B/32, significantly improved scene classification, achieving a top F1 score of 91.1%. These results demonstrate the ability of the system to deliver rapid and precise scene recognition, which can be used to meet the critical requirements of Advanced Driver Assistance Systems (ADAS). This study shows the potential of CLIP models to provide scalable and efficient frameworks for dynamic scene understanding and classification. Furthermore, this work lays the groundwork for advanced autonomous vehicle technologies by fostering a deeper understanding of driver behavior, road conditions, and safety-critical scenarios, marking a significant step toward smarter, safer, and more context-aware autonomous driving systems.

自动驾驶CLIP模型场景理解边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。