arXiv:2409.01014cs.CVcs.AI2024-09ICRA

用预训练扩散模型生成与鸟瞰图对齐的街景图像。

From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model

论文配图:From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model
图 1 · 摘自论文原文
  • 通过神经视图转换建立鸟瞰图与街景的形状对应关系。
  • 利用语义分割图引导微调后的潜空间扩散模型生成图像。
  • 生成图像在视角和风格上保持一致,适合自动驾驶场景建模。

本文研究鸟瞰图(BEV)生成任务,即将鸟瞰图地图转换为对应的多视角街景图像。由于具有统一的空间表示,便于多传感器融合,鸟瞰图在自动驾驶中至关重要。从鸟瞰图准确生成街景图像,有助于刻画复杂交通场景并提升驾驶算法性能。与此同时,基于扩散的条件图像生成模型已展现出卓越效果,能生成多样、高质量且条件对齐的结果。然而,这些模型训练需要大量数据和计算资源。因此,探索如何微调先进模型(如Stable Diffusion)以完成特定条件生成任务成为重要方向。本文提出一个实用框架,用于从鸟瞰图布局生成图像,包含两个核心组件:神经视图转换和街景图像生成。神经视图转换阶段通过学习鸟瞰图与透视视图间的形状对应关系,将鸟瞰图转化为对齐的多视角语义分割图。随后,街景图像生成阶段利用这些分割图作为条件,指导微调后的潜空间扩散模型生成图像。该微调过程确保了视角和风格的一致性。模型充分利用大规模预训练扩散模型在交通场景中的生成能力,有效生成多样化且条件一致的街景图像。

原文摘要 · Abstract (English)

We explore Bird's-Eye View (BEV) generation, converting a BEV map into its corresponding multi-view street images. Valued for its unified spatial representation aiding multi-sensor fusion, BEV is pivotal for various autonomous driving applications. Creating accurate street-view images from BEV maps is essential for portraying complex traffic scenarios and enhancing driving algorithms. Concurrently, diffusion-based conditional image generation models have demonstrated remarkable outcomes, adept at producing diverse, high-quality, and condition-aligned results. Nonetheless, the training of these models demands substantial data and computational resources. Hence, exploring methods to fine-tune these advanced models, like Stable Diffusion, for specific conditional generation tasks emerges as a promising avenue. In this paper, we introduce a practical framework for generating images from a BEV layout. Our approach comprises two main components: the Neural View Transformation and the Street Image Generation. The Neural View Transformation phase converts the BEV map into aligned multi-view semantic segmentation maps by learning the shape correspondence between the BEV and perspective views. Subsequently, the Street Image Generation phase utilizes these segmentations as a condition to guide a fine-tuned latent diffusion model. This finetuning process ensures both view and style consistency. Our model leverages the generative capacity of large pretrained diffusion models within traffic contexts, effectively yielding diverse and condition-coherent street view images.

图像生成自动驾驶扩散模型鸟瞰图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。