arXiv:2505.24260cs.AI2025-05被引 10

用分步生成+多模态模型,让设计师全程掌控城市设计过程。

Generative AI for Urban Design: A Stepwise Approach Integrating Human Expertise with Multimodal Diffusion Models

  • 分三阶段生成:道路用地、建筑布局、细节渲染,贴合真实设计流程。
  • 在芝加哥和纽约数据上,生成方案在准确性、合规性、多样性上均优于基线。
  • 适合城市规划师与设计师,强调人机协作与迭代优化。

城市设计涉及多重约束与多方协作,传统方法效率有限。生成式人工智能(GenAI)有望提升设计效率并促进创意沟通,但多数现有方法脱离人类工作流,采用难以控制的端到端生成方式,忽略实际设计中的迭代特性。本文提出一种分步式生成城市设计框架,将多模态扩散模型与人类专业经验结合,实现更灵活可控的设计过程。该框架分为三个阶段:(1)道路网络与用地规划,(2)建筑布局规划,(3)详细规划与渲染。每个阶段基于文本提示与图像约束生成初步方案,供设计师评审与修改。我们构建评估体系,衡量生成方案的保真度、合规性与多样性。在芝加哥与纽约市数据上的实验表明,本框架在三项指标上均优于基线模型与端到端方法。研究证明,多模态扩散模型与分步生成能有效保持人类控制力,支持迭代优化,为城市设计中的人机协同提供坚实基础。

原文摘要 · Abstract (English)

Urban design is a multifaceted process that demands careful consideration of site-specific constraints and collaboration among diverse professionals and stakeholders. The advent of generative artificial intelligence (GenAI) offers transformative potential by improving the efficiency of design generation and facilitating the communication of design ideas. However, most existing approaches are not well integrated with human design workflows. They often follow end-to-end pipelines with limited control, overlooking the iterative nature of real-world design. This study proposes a stepwise generative urban design framework that integrates multimodal diffusion models with human expertise to enable more adaptive and controllable design processes. Instead of generating design outcomes in a single end-to-end process, the framework divides the process into three key stages aligned with established urban design workflows: (1) road network and land use planning, (2) building layout planning, and (3) detailed planning and rendering. At each stage, multimodal diffusion models generate preliminary designs based on textual prompts and image-based constraints, which can then be reviewed and refined by human designers. We design an evaluation framework to assess the fidelity, compliance, and diversity of the generated designs. Experiments using data from Chicago and New York City demonstrate that our framework outperforms baseline models and end-to-end approaches across all three dimensions. This study underscores the benefits of multimodal diffusion models and stepwise generation in preserving human control and facilitating iterative refinements, laying the groundwork for human-AI interaction in urban design solutions.

城市设计多模态生成扩散模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。