用大模型自动理解交通场景,提升自动驾驶数据标注效率
Scenario Understanding of Traffic Scenes Through Large Visual Language Models
- 用GPT-4、LLaVA等大视觉语言模型自动分析交通场景
- 在自建数据集和BDD100K上实现高效场景分类
- 适合自动驾驶数据标注与跨域泛化研究者使用
自动驾驶的深度学习模型依赖大规模数据实现高性能,但受限于领域特定的数据分布,泛化能力常受影响。手动标注虽有价值,却耗时费力,成为数据标注瓶颈。大视觉语言模型(LVLM)通过上下文查询实现图像分析与分类,无需重新训练即可适应新类别,展现出巨大潜力。本研究评估了GPT-4与LLaVA在自建数据集及BDD100K上的表现,提出一种可扩展的自动标注流水线,结合定量指标与定性分析,验证了LVLM在理解城市交通场景方面的有效性,证明其是推动自动驾驶数据驱动发展的高效工具。
原文摘要 · Abstract (English)
Deep learning models for autonomous driving, encompassing perception, planning, and control, depend on vast datasets to achieve their high performance. However, their generalization often suffers due to domain-specific data distributions, making an effective scene-based categorization of samples necessary to improve their reliability across diverse domains. Manual captioning, though valuable, is both labor-intensive and time-consuming, creating a bottleneck in the data annotation process. Large Visual Language Models (LVLMs) present a compelling solution by automating image analysis and categorization through contextual queries, often without requiring retraining for new categories. In this study, we evaluate the capabilities of LVLMs, including GPT-4 and LLaVA, to understand and classify urban traffic scenes on both an in-house dataset and the BDD100K. We propose a scalable captioning pipeline that integrates state-of-the-art models, enabling a flexible deployment on new datasets. Our analysis, combining quantitative metrics with qualitative insights, demonstrates the effectiveness of LVLMs to understand urban traffic scenarios and highlights their potential as an efficient tool for data-driven advancements in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。