构建开放数据集提升自动驾驶模型应对突发复杂场景能力
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

- 从8个开源数据集提取超8万段视频,构建四类未结构化场景新数据集
- 训练后模型在NeuroNCAP闭环测试中碰撞率下降,nuScenes轨迹预测接近顶尖水平
- 配套问答诊断工具可精准定位感知、预测与规划模块的改进效果
自动驾驶视觉-语言-动作(VLA)模型在非结构化异常场景下表现不佳,主要受限于缺乏针对性评估基准。为此,我们提出Impromptu VLA,核心贡献是构建了一个包含超过80,000段精心筛选视频片段的数据集,这些片段源自8个开源大规模数据集的200万原始片段。该数据集基于全新的四类挑战性非结构化场景分类体系,包含丰富的面向规划的问答标注和动作轨迹。关键实验表明,使用本数据集训练的VLA模型在现有基准上实现显著性能提升:闭环NeuroNCAP评分提高,碰撞率降低,开环nuScenes轨迹预测的L2精度接近当前最优水平。此外,我们的问答评测套件有效揭示了视觉语言模型在感知、预测与规划能力上的进步。代码、数据与模型已公开于https://github.com/ahydchh/Impromptu-VLA。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets. This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories. Crucially, experiments demonstrate that VLAs trained with our dataset achieve substantial performance gains on established benchmarks--improving closed-loop NeuroNCAP scores and collision rates, and reaching near state-of-the-art L2 accuracy in open-loop nuScenes trajectory prediction. Furthermore, our Q&A suite serves as an effective diagnostic, revealing clear VLM improvements in perception, prediction, and planning. Our code, data and models are available at https://github.com/ahydchh/Impromptu-VLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。