arXiv:2606.02956cs.CVcs.LG2026-06被引 1

KITScenes发布高精度多模态数据集,助力自动驾驶感知与地图构建

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

论文配图:The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
图 1 · 摘自论文原文
  • 融合高清摄像头、400米以上长距激光雷达与4D成像雷达,实现多传感器同步采集
  • 首次在公开数据集中完整标注3D交通灯等元素,地图具备拓扑连通性与重投影精度
  • 覆盖复杂城市街景,适合研究长距离感知、端到端驾驶等智能体空间学习任务

现有自动驾驶数据集在传感器精度、地图完整性或地理多样性方面仍有不足。我们提出KITScenes Multimodal,一个基于高保真传感器与地图的欧洲数据集。其全同步传感器套件包括高分辨率全局快门相机、探测距离超过400米的远距激光雷达、4D成像雷达以及冗余的GNSS/INS定位系统。所构建的HD地图据我们所知是当前最完整的传感器数据集地图,经开源软件自动驾驶测试验证。首次在公开数据集中以重投影精度完整映射所有行车相关交通元素(如红绿灯)并实现拓扑连通。数据采集于街道布局不规则、混合交通模式的城市,显著拓展了地理多样性。我们还引入四个基准任务,分别推进具身AI的空间学习:在线高精地图构建、长距离深度估计、新视角合成与端到端驾驶。项目主页:https://kitscenes.com/

原文摘要 · Abstract (English)

Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400m, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, validated through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving. Project page: https://kitscenes.com/

自动驾驶多模态数据集高精地图长距感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。