用卫星图与生成AI预测全球城市未来形态
Envisioning global urban development with satellite imagery and generative AI
- 融合文本提示与地理约束,生成高保真城市图像
- 在500个最大都市区生成多样真实场景,支持可控设计
- 可辅助规划决策,提升碳排放等预测任务性能
城市化是人类历史的决定性力量,但以往研究多将其视为预测任务,未能体现其生成本质。本文设计了一个多模态生成式AI框架,用于在全球尺度上构想可持续城市发展模式。通过整合文本提示与地理空间控制,该框架可生成500个全球最大都市区的高保真、多样化且逼真的城市卫星影像。用户可设定发展目标,生成符合目标的多种场景,外观由文本提示和地理约束共同调控。系统还能学习周边环境特征,支持城市更新实践。此外,模型编码并解析了城市形态的潜在表征,实现跨城市的空间知识迁移,并可提升碳排放预测等下游任务表现。专家评估表明,生成图像与真实图像相当。本研究为加速城市规划提供了新范式,支持全球城市的基于场景的规划流程。
原文摘要 · Abstract (English)
Urban development has been a defining force in human history, shaping cities for centuries. However, past studies mostly analyze such development as predictive tasks, failing to reflect its generative nature. Therefore, this study designs a multimodal generative AI framework to envision sustainable urban development at a global scale. By integrating prompts and geospatial controls, our framework can generate high-fidelity, diverse, and realistic urban satellite imagery across the 500 largest metropolitan areas worldwide. It enables users to specify urban development goals, creating new images that align with them while offering diverse scenarios whose appearance can be controlled with text prompts and geospatial constraints. It also facilitates urban redevelopment practices by learning from the surrounding environment. Beyond visual synthesis, we find that it encodes and interprets latent representations of urban form for global cross-city learning, successfully transferring styles of urban environments across a global spatial network. The latent representations can also enhance downstream prediction tasks such as carbon emission prediction. Further, human expert evaluation confirms that our generated urban images are comparable to real urban images. Overall, this study presents innovative approaches for accelerated urban planning and supports scenario-based planning processes for worldwide cities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。