用AI分析城市空间布局,帮市民推荐合理的小尺度改造方案
Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models

- 基于共现嵌入识别常见空间组合模式
- 可生成5个统计上最可能的搭配对象
- 适合参与式城市设计与公共空间优化
本文提出一种人机协同的计算机视觉框架,利用生成式AI为公共空间提供微观尺度的设计干预建议,并支持更持续、本地化的公众参与。系统采用Grounding DINO和经筛选的ADE20K数据集作为城市建成环境的代理,检测城市物体并构建共现嵌入,揭示常见的空间配置模式。基于此分析,用户可获得五个统计上最可能的锚点对象补全项。随后,视觉语言模型对场景图像及选定组合进行推理,建议第三个能完成更复杂城市策略的物体。整个流程保持用户对选择与优化的控制权,旨在通过扎根于日常模式与生活经验,推动超越自上而下总体规划的参与式设计。
原文摘要 · Abstract (English)
This paper introduces a human-in-the-loop computer vision framework that uses generative AI to propose micro-scale design interventions in public space and support more continuous, local participation. Using Grounding DINO and a curated subset of the ADE20K dataset as a proxy for the urban built environment, the system detects urban objects and builds co-occurrence embeddings that reveal common spatial configurations. From this analysis, the user receives five statistically likely complements to a chosen anchor object. A vision language model then reasons over the scene image and the selected pair to suggest a third object that completes a more complex urban tactic. The workflow keeps people in control of selection and refinement and aims to move beyond top-down master planning by grounding choices in everyday patterns and lived experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。