arXiv:2510.19195cs.CVcs.AI2025-10被引 14

用3D引导生成多视角驾驶视频,提升感知模型对极端场景的识别能力。

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

  • 分解视频为3D引导图,渲染3D资产生成多视角真实视频。
  • 在不同训练轮次下均显著提升下游感知模型性能,最高增益达18.7%。
  • 适合自动驾驶感知模型训练,尤其强化对罕见场景的泛化能力。

近期的驾驶世界模型可生成高质量的RGB或多模态视频。现有方法多关注生成质量和可控性,却忽视了对下游感知任务的评估,而后者对自动驾驶至关重要。传统方法采用先在合成数据上预训练、再在真实数据上微调的策略,训练周期是基线(仅真实数据)的两倍。当基线也增加至双倍轮次时,合成数据的优势几乎消失。为真正验证合成数据的价值,本文提出Dream4Drive——一种专为增强下游感知任务设计的合成数据生成框架。该框架首先将输入视频分解为多个3D感知引导图,然后将3D资产渲染到这些引导图上,并微调驾驶世界模型以生成编辑过的多视角逼真视频,可用于训练下游感知模型。该方法实现了前所未有的大规模多视角极端场景生成能力,显著提升自动驾驶中的边缘场景感知性能。为促进后续研究,我们还发布了大型3D资产数据集DriveObj3D,涵盖典型驾驶场景类别,支持多样化的3D感知视频编辑。通过全面实验验证,Dream4Drive在不同训练轮次下均能有效提升下游感知模型性能。

原文摘要 · Abstract (English)

Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. Existing methods primarily focus on metrics related to generation quality and controllability. However, they often overlook the evaluation of downstream perception tasks, which are $\mathbf{really\ crucial}$ for the performance of autonomous driving. Existing methods usually leverage a training strategy that first pretrains on synthetic data and finetunes on real data, resulting in twice the epochs compared to the baseline (real data only). When we double the epochs in the baseline, the benefit of synthetic data becomes negligible. To thoroughly demonstrate the benefit of synthetic data, we introduce Dream4Drive, a novel synthetic data generation framework designed for enhancing the downstream perception tasks. Dream4Drive first decomposes the input video into several 3D-aware guidance maps and subsequently renders the 3D assets onto these guidance maps. Finally, the driving world model is fine-tuned to produce the edited, multi-view photorealistic videos, which can be used to train the downstream perception models. Dream4Drive enables unprecedented flexibility in generating multi-view corner cases at scale, significantly boosting corner case perception in autonomous driving. To facilitate future research, we also contribute a large-scale 3D asset dataset named DriveObj3D, covering the typical categories in driving scenarios and enabling diverse 3D-aware video editing. We conduct comprehensive experiments to show that Dream4Drive can effectively boost the performance of downstream perception models under various training epochs. Page: https://wm-research.github.io/Dream4Drive/ GitHub Link: https://github.com/wm-research/Dream4Drive

自动驾驶合成数据3D生成感知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。