将普通行车记录仪视频转为自动驾驶所需多传感器数据
Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving

- 用4D高斯点云重建真实车载数据,生成配对训练样本
- 通过扩散模型将单目视频转为多视角图像与激光点云
- 可扩展利用网络公开视频,提升自动驾驶训练多样性
自动驾驶系统(ADS)的鲁棒训练与验证依赖海量多样数据。车企自有车队采集的高保真数据虽质量高,但规模有限,传感器配置、地理分布及长尾行为覆盖不足。相比之下,来自行车记录仪等的野外视频数据量大、场景丰富,能捕捉关键长尾事件和新环境,但其非结构化特性难以适配需多模态输入的自动驾驶系统。为此,我们提出Sensor2Sensor,一种生成式建模范式,可将野外单目行车记录视频转化为包含多视角图像与激光雷达点云的高保真多模态传感器数据。核心挑战是缺乏成对训练数据,我们通过4D高斯点云重建(4DGS)与新视角渲染,将真实车载日志转换为行车记录仪风格视频作为训练对。随后,Sensor2Sensor采用扩散模型实现生成转换。我们在真实数据上进行全面定量评估,验证生成数据的保真度与真实性。实验表明,该方法可成功将复杂互联网及行车记录视频转化为逼真的多模态数据,显著拓展自动驾驶开发可用外部数据源。
原文摘要 · Abstract (English)
Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as geographic and long-tail-behavioral coverage. In contrast, in-the-wild data from sources like dashcams offers immense scale and diversity, capturing critical long-tail scenarios and novel environments. However, this unstructured, in-the-wild video data is incompatible with ADS expecting structured, multi-modal sensor inputs for validation and training. To bridge this data gap, we propose Sensor2Sensor, a novel generative modeling paradigm that translates in-the-wild monocular dashcam videos into a high-fidelity, multi-modal sensor suite (AV logs) comprising multi-view camera images and LiDAR point clouds. A core challenge is the lack of paired training data. We address this by converting real AV logs into dashcam-style videos via 4D Gaussian Splatting (4DGS) reconstruction and novel-view rendering. Sensor2Sensor then utilizes a diffusion architecture to perform the generative conversion. We perform comprehensive quantitative evaluations on the fidelity and realism of the generated sensor data. We demonstrate Sensor2Sensor's practical utility by converting challenging in-the-wild internet and dashcam footage into realistic, multi-modal data formats, further unlocking vast external data sources for AV development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。