arXiv:2601.19582cs.CV2026-01

构建首个大规模第一人称驾驶数据集,用于评估视觉语言模型在安全感知中的表现。

ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving

  • 从公开视频构建3847小时第一人称驾驶数据,覆盖63国1210城
  • 提出多维度基准评测,发现现有模型在几何感知上仍有明显不足
  • 适用于自动驾驶安全评估与跨区域鲁棒性研究的开发者

本文介绍ScenePilot-4K,一个面向自动驾驶中安全感知的大型第一人称视觉语言学习与评估数据集。该数据集基于公开在线驾驶视频构建,包含3,847小时视频和2770万张前视图像,覆盖63个国家/地区和1,210个城市。通过统一的多阶段标注流程,提供场景级自然语言描述、风险评估标签、关键参与者标注、自车轨迹及相机参数。基于此数据集,我们建立ScenePilot-Bench标准化基准,从场景理解、空间感知、运动规划和GPT语义对齐四个互补维度评估视觉语言模型。基准包含细粒度指标和跨区域、跨交通域的泛化设置,可暴露模型在域偏移下的鲁棒性。在主流开源与专有视觉语言模型上的基线结果显示,当前模型在高层语义层面仍具竞争力,但在几何感知与规划推理方面仍存在显著局限。除数据集外,所提出的标注流程可作为从公开互联网驾驶视频规模化构建数据集的可复用范式。代码与补充材料见:https://github.com/yjwangtj/ScenePilot-4K,数据集可在 https://huggingface.co/datasets/larswangtj/ScenePilot-4K 获取。

原文摘要 · Abstract (English)

In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline. Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts. Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning. Beyond the released dataset itself, the proposed annotation pipeline serves as a reusable and extensible recipe for scalable dataset construction from public Internet driving videos. The codes and supplementary materials are available at: https://github.com/yjwangtj/ScenePilot-4K, with the dataset available at https://huggingface.co/datasets/larswangtj/ScenePilot-4K.

自动驾驶视觉语言数据集安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。