用多模态信息生成更真实的城市形态,突破单纯几何模仿的局限。
From Geometric Mimicry to Comprehensive Generation: A Context-Informed Multimodal Diffusion Model for Urban Morphology Synthesis
- 融合图像、文本、元数据与建筑轮廓,多模态协同控制生成过程。
- 相比单模态方法,形态保真度提升71.01%(FID降至50.94),空间重叠率提高38.46%(达0.36)。
- 支持跨城市风格迁移与未知城市零样本生成,适合城市规划与数字孪生应用。
城市形态决定城市功能与活力。现有模拟方法常将形态生成简化为几何问题,缺乏对语义与地理上下文的深层理解。为此,本文提出ControlCity,一种通过多模态信息融合实现综合城市形态生成的扩散模型。构建了包含22座全球城市“图像-文本-元数据-建筑轮廓”的四维数据集。ControlCity利用多维信息作为联合控制条件:增强的ControlNet从图像中编码空间约束,文本提供语义引导,元数据赋予地理先验,共同指导生成过程。实验表明,相比单模态基线,该方法在形态保真度上显著提升,视觉误差(FID)降低71.01%至50.94,空间重叠率(MIoU)提升38.46%至0.36。模型还展现出强泛化能力与可控性,可实现跨城市风格迁移与未知城市零样本生成。消融实验揭示图像、文本与元数据在生成中的差异化作用。研究证实,多模态融合是实现从‘几何模仿’到‘理解驱动的综合生成’的关键,为城市形态研究与应用提供新范式。
原文摘要 · Abstract (English)
Urban morphology is fundamental to determining urban functionality and vitality. Prevailing simulation methods, however, often oversimplify morphological generation as a geometric problem, lacking a profound understanding of urban semantics and geographical context. To address this limitation, this study proposes ControlCity, a diffusion model that achieves comprehensive urban morphology generation through multimodal information fusion. We first constructed a quadruple dataset comprising ``image-text-metadata-building footprints" from 22 cities worldwide. ControlCity utilizes these multidimensional information as joint control conditions, where an enhanced ControlNet architecture encodes spatial constraints from images, while text and metadata provide semantic guidance and geographical priors respectively, collectively directing the generation process. Experimental results demonstrate that compared to unimodal baselines, this method achieves significant advantages in morphological fidelity, with visual error (FID) reduced by 71.01%, reaching 50.94, and spatial overlap (MIoU) improved by 38.46%, reaching 0.36. Furthermore, the model demonstrates robust knowledge generalization and controllability, enabling cross-city style transfer and zero-shot generation for unknown cities. Ablation studies further reveal the distinct roles of images, text, and metadata in the generation process. This study confirms that multimodal fusion is crucial for achieving the transition from ``geometric mimicry" to ``understanding-based comprehensive generation," providing a novel paradigm for urban morphology research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。