arXiv:2512.22706cs.CV2025-12

统一生成真实3D车辆插入与新视角图像,提升自动驾驶仿真数据多样性。

SCPainter: A Unified Framework for Realistic 3D Asset Insertion and Novel View Synthesis

  • 用3D高斯溅射+扩散模型联合生成车辆与场景新视角
  • 在Waymo数据集上实现高质量车辆插入与视角合成
  • 适合自动驾驶数据增强与长尾场景生成

3D资产插入与新视角合成(NVS)是自动驾驶仿真中的关键环节,有助于提升训练数据的多样性和覆盖度,包括罕见驾驶场景,从而增强模型鲁棒性与安全性。现有方法分别处理资产插入与视角合成,缺乏联动。为此,我们提出SCPainter(Street Car Painter),融合3D高斯溅射(GS)车辆表示与3D场景点云,结合扩散模型,统一实现真实感3D资产插入与新视角图像生成。将3D GS资产与场景点云共同投影至新视角,作为扩散模型的条件,生成高质量图像。在Waymo Open Dataset上的评估表明,该框架可有效支持3D资产插入与多视角合成,推动多样化、真实感驾驶数据的创建。

原文摘要 · Abstract (English)

3D Asset insertion and novel view synthesis (NVS) are key components for autonomous driving simulation, enhancing the diversity of training data. With better training data that is diverse and covers a wide range of situations, including long-tailed driving scenarios, autonomous driving models can become more robust and safer. This motivates a unified simulation framework that can jointly handle realistic integration of inserted 3D assets and NVS. Recent 3D asset reconstruction methods enable reconstruction of dynamic actors from video, supporting their re-insertion into simulated driving scenes. While the overall structure and appearance can be accurate, it still struggles to capture the realism of 3D assets through lighting or shadows, particularly when inserted into scenes. In parallel, recent advances in NVS methods have demonstrated promising results in synthesizing viewpoints beyond the originally recorded trajectories. However, existing approaches largely treat asset insertion and NVS capabilities in isolation. To allow for interaction with the rest of the scene and to enable more diverse creation of new scenarios for training, realistic 3D asset insertion should be combined with NVS. To address this, we present SCPainter (Street Car Painter), a unified framework which integrates 3D Gaussian Splat (GS) car asset representations and 3D scene point clouds with diffusion-based generation to jointly enable realistic 3D asset insertion and NVS. The 3D GS assets and 3D scene point clouds are projected together into novel views, and these projections are used to condition a diffusion model to generate high quality images. Evaluation on the Waymo Open Dataset demonstrate the capability of our framework to enable 3D asset insertion and NVS, facilitating the creation of diverse and realistic driving data.

3D生成自动驾驶扩散模型新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。