让AI看懂鸟瞰图地图,实现交通场景的智能问答与生成
ChatBEV: A Visual Language Model that Understands BEV Maps
- 构建13.7万条问题的BEV地图理解基准,覆盖全局与车辆交互任务
- 训练专用视觉语言模型ChatBEV,可精准回答复杂交通场景问题
- 支持文本驱动生成真实交通场景,适合自动驾驶研发人员使用
交通场景理解对智能交通系统和自动驾驶至关重要,确保车辆安全高效运行。尽管视觉语言模型(VLMs)在整体场景理解方面展现潜力,但其在交通场景尤其是鸟瞰图(BEV)地图上的应用仍不充分。现有方法常受限于任务设计狭窄和数据量不足,难以实现全面理解。为此,我们提出ChatBEV-QA,一个包含超过13.7万道问题的新颖BEV视觉问答基准,涵盖全局场景理解、车-车道交互及车-车交互等多样化任务。该基准通过新型数据采集流水线生成可扩展且信息丰富的VQA数据。我们进一步微调专用视觉语言模型ChatBEV,使其能理解多种问题提示并从BEV地图中提取上下文相关的信息。此外,我们提出一种语言驱动的交通场景生成流程,利用ChatBEV实现地图理解与文本对齐的导航指导,显著提升真实且一致的交通场景生成效果。相关数据集、代码与微调模型将公开发布。
原文摘要 · Abstract (English)
Traffic scene understanding is essential for intelligent transportation systems and autonomous driving, ensuring safe and efficient vehicle operation. While recent advancements in VLMs have shown promise for holistic scene understanding, the application of VLMs to traffic scenarios, particularly using BEV maps, remains under explored. Existing methods often suffer from limited task design and narrow data amount, hindering comprehensive scene understanding. To address these challenges, we introduce ChatBEV-QA, a novel BEV VQA benchmark contains over 137k questions, designed to encompass a wide range of scene understanding tasks, including global scene understanding, vehicle-lane interactions, and vehicle-vehicle interactions. This benchmark is constructed using an novel data collection pipeline that generates scalable and informative VQA data for BEV maps. We further fine-tune a specialized vision-language model ChatBEV, enabling it to interpret diverse question prompts and extract relevant context-aware information from BEV maps. Additionally, we propose a language-driven traffic scene generation pipeline, where ChatBEV facilitates map understanding and text-aligned navigation guidance, significantly enhancing the generation of realistic and consistent traffic scenarios. The dataset, code and the fine-tuned model will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。