arXiv:2507.06971cs.CVcs.RO2025-07中稿 · ICRA

用扩散模型生成360度街景,还能精准控制画面内容。

Hallucinating 360°: Panoramic Street-View Generation via Local Scenes Diffusion and Probabilistic Prompting

  • 通过局部场景扩散重建全景图像,弥补针孔采样损失
  • 生成图像在无参考质量评估中优于原始拼接图
  • 适合自动驾驶数据增强与可控图像生成研究者

全景感知对自动驾驶具有重要意义,可实现单次拍摄获取360°全向视图。然而,自动驾驶是数据驱动任务,完整全景数据采集需复杂采样系统与标注流程,耗时且人力密集。现有街景生成模型虽具较强再生能力,但仅能学习现有数据集的固定分布,无法利用拼接针孔图像作为监督信号。本文提出首个面向自动驾驶的全景生成方法Percep360,支持基于拼接全景数据的连贯性与可控性生成。Percep360聚焦一致性与可控性:为克服针孔采样导致的信息损失,提出局部场景扩散法(LSDM),将全景生成重构为连续空间扩散过程,弥合不同数据分布差异;为实现可控生成,提出概率提示法(PPM),动态选择最相关控制线索。从图像质量(含无参考与有参考评估)、可控性及真实世界鸟瞰图(BEV)分割实用性三方面评估,生成数据在无参考质量指标上持续优于原始拼接图像,并提升下游感知模型性能。代码将公开于https://github.com/FeiT-FeiTeng/Percep360。

原文摘要 · Abstract (English)

Panoramic perception holds significant potential for autonomous driving, enabling vehicles to acquire a comprehensive 360° surround view in a single shot. However, autonomous driving is a data-driven task. Complete panoramic data acquisition requires complex sampling systems and annotation pipelines, which are time-consuming and labor-intensive. Although existing street view generation models have demonstrated strong data regeneration capabilities, they can only learn from the fixed data distribution of existing datasets and cannot leverage stitched pinhole images as a supervisory signal. In this paper, we propose the first panoramic generation method Percep360 for autonomous driving. Percep360 enables coherent generation of panoramic data with control signals based on the stitched panoramic data. Percep360 focuses on two key aspects: coherence and controllability. Specifically, to overcome the inherent information loss caused by the pinhole sampling process, we propose the Local Scenes Diffusion Method (LSDM). LSDM reformulates the panorama generation as a spatially continuous diffusion process, bridging the gaps between different data distributions. Additionally, to achieve the controllable generation of panoramic images, we propose a Probabilistic Prompting Method (PPM). PPM dynamically selects the most relevant control cues, enabling controllable panoramic image generation. We evaluate the effectiveness of the generated images from three perspectives: image quality assessment (i.e., no-reference and with reference), controllability, and their utility in real-world Bird's Eye View (BEV) segmentation. Notably, the generated data consistently outperforms the original stitched images in no-reference quality metrics and enhances downstream perception models. The source code will be publicly available at https://github.com/FeiT-FeiTeng/Percep360.

全景生成扩散模型自动驾驶可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。