arXiv:2606.07366cs.CVcs.LG2026-06

将真实行车记录仪视频转为可模拟的4D驾驶数据,拓展长尾场景研究。

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos

论文配图:Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos
图 1 · 摘自论文原文
  • 从单目行车视频重建带地理坐标的4D驾驶场景
  • 构建涵盖17城2.7M物体的ROADWork4D数据集,验证2201个闭环场景
  • 提升视角合成质量19%,适合自动驾驶规划与感知研究

自动驾驶仿真通常依赖少数城市采集的数据或人工构建的合成场景。行车记录仪视频覆盖更广的地理和交通情境,包括罕见的长尾场景(如施工区),但因难以从单目真实视频中恢复精确4D场景,长期难用于仿真。本文提出Dash2Sim框架,将真实单目行车视频转化为符合现有仿真器要求的度量级、地理参考4D驾驶日志,并通过独立维护的地图进行无标注验证。我们基于大规模视频语料构建了ROADWork4D基准数据集,包含4,244个场景、270万3D物体,覆盖17个城市。在验证子集ROADWork4D-CL(2,201个场景)上评估闭环规划器发现:规则与混合规划器比学习型更好,但仍无法完成临时施工区所需的变道动作。此外,由Dash2Sim恢复的稠密深度使新视角合成在感知指标上提升最高达19%,表明其可为单目视频驱动的闭环传感器仿真提供丰富条件。

原文摘要 · Abstract (English)

Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations and situations, including rare or long-tailed scenarios. They are considered less usable for simulation because it is difficult to recover accurate 4D scenes from monocular in-the-wild videos. Work zones are one such class of long-tailed situations that dashcams capture. We present Dash2Sim, a framework that turns in-the-wild monocular dashcam videos into metric, geo-referenced 4D driving logs compatible with existing simulators, and verifies eachone against an independently maintained map without annotations. We apply Dash2Sim to a large video corpus to create the ROADWork4D benchmark dataset, which spans 4,244 scenes with 2.7M 3D objects across 17 cities. On a verified subset ROADWork4D-CL (2,201 scenes), we study privileged closed-loop planners and find that work zone scenarios are difficult: while rule-based and hybrid planners generalize better than learning-based ones, all fall short, failing to make the lane changes that temporary work zone channels require. Beyond planning, dense depth recovered by Dash2Sim improves novel-view synthesis quality by up to 19% on perceptual metrics, suggesting its potential to provide rich conditioning for closed-loop sensor simulation from monocular videos.

自动驾驶视频仿真4D重建长尾场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。