为自动驾驶交通行为理解构建多模态模型评测基准与训练数据
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
- 设计TB-Bench评测集,覆盖8个交通行为感知任务
- 微调后模型平均准确率达85%,远超原始模型不足35%的表现
- 数据可迁移,提升其他交通数据集上的性能,适合自动驾驶研究者
多模态大语言模型(MLLMs)在自动驾驶中的应用受限于缺乏交通领域专用训练数据及缺乏时空理解的评测基准。本文提出TB-Bench,一个面向驾驶视角的综合性评测基准,涵盖8个感知任务。同时构建了两个视觉-语言指令微调数据集:TB-100k和TB-250k,以及简单有效的基线模型。大量实验表明,现有MLLMs表现不佳,即使强大模型GPT-4o在任务中平均准确率也低于35%。而使用TB-100k或TB-250k微调后,基线模型平均准确率可达85%。此外,将TB-100k与其他交通数据集联合训练,可提升后者性能。本工作为推动MLLM在自动驾驶感知、预测与规划环节的应用提供了完整基准、高质量数据与有效基线。
原文摘要 · Abstract (English)
The application of Multi-modal Large Language Models (MLLMs) in Autonomous Driving (AD) faces significant challenges due to their limited training on traffic-specific data and the absence of dedicated benchmarks for spatiotemporal understanding. This study addresses these issues by proposing TB-Bench, a comprehensive benchmark designed to evaluate MLLMs on understanding traffic behaviors across eight perception tasks from ego-centric views. We also introduce vision-language instruction tuning datasets, TB-100k and TB-250k, along with simple yet effective baselines for the tasks. Through extensive experiments, we show that existing MLLMs underperform in these tasks, with even a powerful model like GPT-4o achieving less than 35% accuracy on average. In contrast, when fine-tuned with TB-100k or TB-250k, our baseline models achieve average accuracy up to 85%, significantly enhancing performance on the tasks. Additionally, we demonstrate performance transfer by co-training TB-100k with another traffic dataset, leading to improved performance on the latter. Overall, this study represents a step forward by introducing a comprehensive benchmark, high-quality datasets, and baselines, thus supporting the gradual integration of MLLMs into the perception, prediction, and planning stages of AD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。