融合视觉语言模型,让城市街道评估既准又可解释。
Interpretable Multimodal Framework for Human-Centered Street Assessment: Integrating Visual-Language Models for Perceptual Urban Diagnostics
- 用视觉-语言大模型融合街景与文本,实现双输出评估。
- 在1.5万张哈尔滨街景图上达到89.3%居民感知一致性。
- 能生成带理由的分析,适合关注人本设计的规划者。
尽管基于影像或地理信息系统(GIS)的客观街道指标已成为城市分析标准,但仍难以捕捉包容性城市设计所必需的主观感知。本研究提出一种新型多模态街道评估框架(MSEF),将视觉变换器(VisualGLM-6B)与大语言模型(GPT-4)融合,实现可解释的双重输出评估。基于中国哈尔滨超过1.5万张标注街景图像,采用LoRA和P-Tuning v2进行参数高效微调。模型在客观特征上达到F1分数0.84,在分层社会经济地理区域中与居民感知聚合结果达成89.3%的一致性。该框架不仅能识别非线性、语义依赖的感知模式(如建筑透明度在住宅区与商业区的不同影响),还揭示了非正式商业提升活力但降低行人舒适度等情境矛盾。通过生成基于注意力机制的自然语言推理,框架实现了感官数据与社会情感推断的桥梁,支持符合可持续发展目标11(宜居城市)的透明诊断,为城市感知建模提供方法创新,并助力规划系统平衡基础设施精度与真实生活体验。
原文摘要 · Abstract (English)
While objective street metrics derived from imagery or GIS have become standard in urban analytics, they remain insufficient to capture subjective perceptions essential to inclusive urban design. This study introduces a novel Multimodal Street Evaluation Framework (MSEF) that fuses a vision transformer (VisualGLM-6B) with a large language model (GPT-4), enabling interpretable dual-output assessment of streetscapes. Leveraging over 15,000 annotated street-view images from Harbin, China, we fine-tune the framework using LoRA and P-Tuning v2 for parameter-efficient adaptation. The model achieves an F1 score of 0.84 on objective features and 89.3 percent agreement with aggregated resident perceptions, validated across stratified socioeconomic geographies. Beyond classification accuracy, MSEF captures context-dependent contradictions: for instance, informal commerce boosts perceived vibrancy while simultaneously reducing pedestrian comfort. It also identifies nonlinear and semantically contingent patterns -- such as the divergent perceptual effects of architectural transparency across residential and commercial zones -- revealing the limits of universal spatial heuristics. By generating natural-language rationales grounded in attention mechanisms, the framework bridges sensory data with socio-affective inference, enabling transparent diagnostics aligned with SDG 11. This work offers both methodological innovation in urban perception modeling and practical utility for planning systems seeking to reconcile infrastructural precision with lived experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。