arXiv:2604.23010cs.CVcs.RO2026-04CVPR被引 11

用隐空间扩散模型生成逼真3D交通物体,支持真实场景下多样化仿真。

GenAssets: Generating in-the-wild 3D Assets in Latent Space

论文配图:GenAssets: Generating in-the-wild 3D Assets in Latent Space
图 1 · 摘自论文原文
  • 先重建后生成:用多场景感知的神经渲染构建高质量隐空间。
  • 在真实驾驶数据上训练,生成完整几何与外观的3D资产。
  • 适合自动驾驶仿真,提升场景多样性和生成效率。

用于交通参与者的高质量3D资产对多传感器仿真至关重要,是实现自动驾驶端到端安全开发的基础。从真实世界数据中构建资产能提升多样性与真实性,但现有基于神经渲染的重建方法速度慢,且生成资产仅在接近原始视角时渲染效果好,限制了其在仿真中的应用。近期基于扩散模型的生成方法虽能构建完整多样资产,但在真实驾驶场景中表现不佳,因观测对象常处于稀疏、有限视场且部分遮挡。本文提出一种3D隐空间扩散模型,利用传感器平台采集的真实世界激光雷达与摄像头数据,学习并生成具有完整几何与外观的高质量3D资产。核心是“先重建后生成”策略:首先在多场景上训练感知遮挡的神经渲染,构建高质量隐空间;再在此隐空间上训练扩散模型。实验表明,该方法优于现有重建与生成方法,实现了仿真中多样化、可扩展的内容生成。

原文摘要 · Abstract (English)

High-quality 3D assets for traffic participants are critical for multi-sensor simulation, which is essential for the safe end-to-end development of autonomy. Building assets from in-the-wild data is key for diversity and realism, but existing neural-rendering based reconstruction methods are slow and generate assets that render well only from viewpoints close to the original observations, limiting their usefulness in simulation. Recent diffusion-based generative models build complete and diverse assets, but perform poorly on in-the-wild driving scenes, where observed actors are captured under sparse and limited fields of view, and are partially occluded. In this work, we propose a 3D latent diffusion model that learns on in-the-wild LiDAR and camera data captured by a sensor platform and generates high-quality 3D assets with complete geometry and appearance. Key to our method is a "reconstruct-then-generate" approach that first leverages occlusion-aware neural rendering trained over multiple scenes to build a high-quality latent space for objects, and then trains a diffusion model that operates on the latent space. We show our method outperforms existing reconstruction and generation based methods, unlocking diverse and scalable content creation for simulation.

3D生成自动驾驶扩散模型隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。