用可控制的地点信息生成真实街景,提升自动驾驶场景识别能力
DiffPlace: Street View Generation via Place-Controllable Diffusion Model Enhancing Place Recognition
- 引入地点ID控制器,让模型根据地点生成一致背景
- 在多个数据集上生成图像质量优于现有方法
- 适合自动驾驶中的场景生成与识别任务
生成模型在逼真图像合成方面取得显著进展,扩散模型在质量和稳定性上表现优异。尽管多视角扩散模型提升了三维感知的街景生成能力,但它们在仅通过文本、鸟瞰图(BEV)地图和物体边界框生成具有地点感知且背景一致的城市场景时仍存在困难,限制了其在场景识别任务中的应用。为此,我们提出DiffPlace,一种引入地点ID控制器的新框架,实现多视角图像的地点可控生成。该控制器采用线性投影、Perceiver Transformer和对比学习,将地点嵌入映射至固定CLIP空间,使模型能在保持背景建筑一致性的同时灵活调整前景物体和天气条件。大量实验,包括定量对比和增强训练评估,表明DiffPlace在生成质量和对视觉场景识别的训练支持方面均优于现有方法。结果凸显了生成模型在场景级与地点感知合成中的潜力,为提升自动驾驶中的场景识别提供了有效路径。
原文摘要 · Abstract (English)
Generative models have advanced significantly in realistic image synthesis, with diffusion models excelling in quality and stability. Recent multi-view diffusion models improve 3D-aware street view generation, but they struggle to produce place-aware and background-consistent urban scenes from text, BEV maps, and object bounding boxes. This limits their effectiveness in generating realistic samples for place recognition tasks. To address these challenges, we propose DiffPlace, a novel framework that introduces a place-ID controller to enable place-controllable multi-view image generation. The place-ID controller employs linear projection, perceiver transformer, and contrastive learning to map place-ID embeddings into a fixed CLIP space, allowing the model to synthesize images with consistent background buildings while flexibly modifying foreground objects and weather conditions. Extensive experiments, including quantitative comparisons and augmented training evaluations, demonstrate that DiffPlace outperforms existing methods in both generation quality and training support for visual place recognition. Our results highlight the potential of generative models in enhancing scene-level and place-aware synthesis, providing a valuable approach for improving place recognition in autonomous driving
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。