arXiv:2605.17527cs.CV2026-05

用扩散模型根据视觉指标生成逼真城市街景,支持规划方案可视化探索。

Designing streetscapes from street-view imagery using diffusion models

论文配图:Designing streetscapes from street-view imagery using diffusion models
图 1 · 摘自论文原文
  • 基于文本与图像双重控制的扩散模型生成街景,实现精准语义一致。
  • 视觉控制使语义一致性提升23.7%~46.4%,建筑视野指标提升超100%。
  • 适合城市规划、设计人员进行可调控的虚拟场景推演与评估。

街景影像(SVI)广泛用于量化城市环境的关键指标,如绿化率、天空可见度或道路视域指数。然而,现有研究多聚焦于当前街景的测量,极少支持替代性或不存在的城市场景生成,而这是城市规划与设计中的核心任务。为此,我们提出一种生成式多模态AI框架,可根据目标视觉指标合成替代街景,实现城市场景的直接视觉探索。我们首先构建了一个多模态数据集,将芝加哥和奥兰多的街景影像与文本描述、分割图、道路掩码及视觉元素定量指标对齐。利用该数据集,我们证明扩散模型可在响应文本与图像控制的同时生成真实且语义一致的街景图像。定量评估显示,引入视觉控制可使语义一致性提升约6%(LPIPS降低),全局视觉真实感保持不变;在奥兰多整体语义一致性提升23.7%,芝加哥提升46.4%(mIoU衡量),建筑视域指标甚至出现超100%的类间提升。街景生成可通过文本与视觉提示进行细粒度控制,当两者冲突时,图像控制始终占主导,体现明确的控制层级。本工作为基于街景影像与扩散模型的街景生成建立了重要基准,并展示了生成式AI在城市情景探索中作为可控制、可扩展、实用化工具的潜力。

原文摘要 · Abstract (English)

Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existing studies largely focus on measuring current streetscapes and rarely support the generation of alternative and non-existing urban scenarios, which is a core task in geospatial disciplines such as urban planning and design. To address this gap, we propose a gener- ative multimodal AI framework that synthesizes alternative streetscapes conditioned on targeted visual metrics, enabling direct visual exploration of urban scenarios. We first construct a multimodal dataset that aligns SVIs with textual descriptions, segmentation maps, road masks, and quantitative metrics of visual elements in Chicago and Orlando. Using this dataset, we demonstrate that diffusion models can produce realistic and semantically consistent streetscape imagery while responding to both textual and imagery controls. Our quantitative evaluations show that incorporating visual controls can improve semantic consistency, reducing the LPIPS index by approximately 6% while maintaining global visual realism. In addition, overall semantic consistency increases by 23.7% in Orlando and 46.4% in Chicago, as measured by the mIoU index, with class-wise gains exceeding even 100% improvement for building view indices. Streetscape generation can be controlled in a fine-grained manner by both visual and textual prompts, and when textual and visual controls conflict, imagery controls consistently dominate, indicating a clear control hierarchy and the importance of further developing visual controls for urban scene generation. Overall, this work establishes an important benchmark for streetscape generation us- ing SVIs and diffusion models, and illustrates how generative AI can serve as a practical, scalable, and controllable approach for urban scenario exploration.

街景生成扩散模型城市规划可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。