arXiv:2506.19488cs.CV2025-06CVPR被引 11

让自动驾驶仿真场景可精准编辑,支持多视角3D一致变化。

SceneCrafter: Controllable Multi-View Driving Scene Editing

  • 基于多视角扩散模型,实现天气、时间等条件的可控编辑
  • 通过合成配对数据提升编辑后场景的几何一致性与真实感
  • 适合自动驾驶系统开发与测试人员使用

仿真对自动驾驶系统开发与评估至关重要。现有方法依赖生成模型合成高度逼真的图像,但缺乏现实基础,难以令人信服。而编辑模型则利用真实驾驶日志中的源场景,可模拟不同交通布局、行为及运行条件(如天气、时段)。然而,驾驶仿真中的图像编辑面临新挑战:(1) 跨摄像头3D一致性需求;(2) 从含遮挡的真实数据中学习“空街”先验;(3) 生成带编辑差异但保持布局与几何一致的图像对。为此,我们提出SceneCrafter,一种用于多相机捕捉驾驶场景的通用编辑框架。基于多视角扩散模型,实现天气、时间、车辆框、高精地图等多模态条件的完全可控编辑。为监督编辑模型,我们提出在Prompt-to-Prompt基础上的新框架,生成全局编辑下的几何一致合成图像对。此外,引入α混合框架,结合基于空街先验训练的掩码训练与多视角重绘范式,实现局部编辑。SceneCrafter在真实感、可控性、3D一致性与编辑质量上均优于现有基线。

原文摘要 · Abstract (English)

Simulation is crucial for developing and evaluating autonomous vehicle (AV) systems. Recent literature builds on a new generation of generative models to synthesize highly realistic images for full-stack simulation. However, purely synthetically generated scenes are not grounded in reality and have difficulty in inspiring confidence in the relevance of its outcomes. Editing models, on the other hand, leverage source scenes from real driving logs, and enable the simulation of different traffic layouts, behaviors, and operating conditions such as weather and time of day. While image editing is an established topic in computer vision, it presents fresh sets of challenges in driving simulation: (1) the need for cross-camera 3D consistency, (2) learning ``empty street" priors from driving data with foreground occlusions, and (3) obtaining paired image tuples of varied editing conditions while preserving consistent layout and geometry. To address these challenges, we propose SceneCrafter, a versatile editor for realistic 3D-consistent manipulation of driving scenes captured from multiple cameras. We build on recent advancements in multi-view diffusion models, using a fully controllable framework that scales seamlessly to multi-modality conditions like weather, time of day, agent boxes and high-definition maps. To generate paired data for supervising the editing model, we propose a novel framework on top of Prompt-to-Prompt to generate geometrically consistent synthetic paired data with global edits. We also introduce an alpha-blending framework to synthesize data with local edits, leveraging a model trained on empty street priors through novel masked training and multi-view repaint paradigm. SceneCrafter demonstrates powerful editing capabilities and achieves state-of-the-art realism, controllability, 3D consistency, and scene editing quality compared to existing baselines.

自动驾驶场景编辑多视角生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。