arXiv:2608.20026cs.CVcs.LG2026-08

用AI分析街景图,快速评估郊区步行友好度

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

  • 用视觉语言模型分析谷歌街景,自动识别街道质量
  • 发现仅局部区域具备良好人行道和绿化,山地住宅区严重不足
  • 适合城市规划者做大规模街道评估,无需实地调研

街道景观质量已成为当代城市规划的核心议题,尤其在倡导步行友好的15分钟城市理念下,可达性与公共空间品质日益成为衡量城市效能的关键。然而,传统实地调查在大范围郊区及城乡结合地带面临时间与资源瓶颈。本文基于SAGAI(Streetscape Analysis with Generative AI)最新版本,在法国尼斯东北外围地区开展面向规划的街道景观质量评估。该开源工作流利用视觉语言模型(VLMs)从谷歌街景图像中实现大规模街道景观分析。新版本改进了图像采集、地理一致视图生成、支持多VLM架构、共识推理及集成分析环境。针对数以千计的街景观测点,评估了人行道存在性、行人出入口密度与植被覆盖率等指标。结果表明,理想街道景观仅存在于部分紧凑开发片区与传统郊区街区,而住宅山坡区域则显著缺失。研究验证了现代VLM在难以实地调研的大范围郊区中支持城市诊断的潜力。论文进一步说明,视觉语言模型的进步可推动基于证据的城市规划,实现可扩展、灵活且可解释的公共空间质量评估。

原文摘要 · Abstract (English)

Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance. However, assessing streetscape qualities across large suburban and peri-urban territories remains challenging due to the time and resource demands of conventional field surveys. This paper presents a planning-oriented assessment of streetscape qualities in the north-eastern periphery of Nice (France) using the latest release of SAGAI (Streetscape Analysis with Generative AI), an open-source workflow that leverages vision-language models (VLMs) for large-scale streetscape analysis from Google Street View imagery. The new release addresses limitations of the original framework through improved image acquisition, geographically consistent view generation, support for multiple VLM architectures, consensus-based inference, and an integrated analytical environment. The workflow is applied to several thousand street-level observations to evaluate qualities relevant to pedestrian-friendly urban environments: sidewalk presence, pedestrian entrance density, and vegetation. The resulting maps reveal that the desired streetscape qualities characterize only a fraction of today's suburban streetscapes, mainly in compact developments and traditional suburban faubourgs, while they are particularly lacking on residential hills. The analysis demonstrates the potential of contemporary VLMs to support urban diagnostics in extensive suburban territories where fieldwork would be prohibitively time-consuming. Beyond the case study, the paper illustrates how recent advances in vision-language models can contribute to evidence-based planning by enabling scalable, flexible, and interpretable assessments of urban public-space quality.

城市规划视觉语言模型街景分析15分钟城市

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。