用视觉语言模型分析城市地标可见性,替代传统3D模拟。
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
- 用街景图像+视觉语言模型检测地标可见性,无需精确3D数据。
- 整体检测准确率87%,伦敦泰晤士河案例中桥梁贡献31%视觉连接。
- 构建可视连通图,适合城市规划与遗产保护场景。
城市可见性分析传统依赖视线(LoS)模拟,但需高精度3D数据且无法反映地标在真实街景中的视觉显著性。本文将地标可见性评估重构为基于图像空间的城市视觉搜索问题,利用广泛存在的街景影像(SVI)。给定目标地标参考图像,通过视觉语言模型(VLM)在方向与缩放可控的街景中检测该地标,成功检测即表示机器识别的可见性。超越单个视角,构建异构可见性图,表征地标、街景位置及中介城市空间间的视觉连通性。该图可揭示视觉连接发生的位置、强度,以及多地标通过共享视觉走廊实现联合连接。在六个全球知名地标结构中,图像方法整体检测准确率达87%,地标可见位置的精确率为68%。在伦敦泰晤士河的第二案例中,可见性图揭示了多地标连接,并识别关键中介位置,其中桥梁约占所有连接的31%。该方法补充了基于LoS的可见性分析,在数据受限环境下提供实用替代方案,也展示了揭示城市环境中视觉对象普遍关联的可能,为城市规划与遗产保护开辟新视角。
原文摘要 · Abstract (English)
Visibility analysis in urban planning has traditionally relied on line-of-sight (LoS) simulations, which capture geometric occlusion. However, these approaches depend on accurate 3D data that is often unavailable and may not adequately represent how visually distinctive urban landmarks are encountered in real streetscapes. We reformulate landmark visibility assessment as an urban visual search problem in image space by leveraging the widespread availability of street view imagery (SVI). Given a reference image of a target landmark, a Vision Language Model (VLM) is applied to detect the landmark in direction- and zoom-controlled SVI. A successful detection indicates machine-recognised landmark visibility at the corresponding viewpoint. Beyond isolated viewpoints, we construct a heterogeneous visibility graph to represent visual connectivity among landmarks, street-view locations, and the urban spaces that mediate them. This graph enables us to map where visual connections occur, how strong they are, and how multiple landmarks become jointly connected through shared visual corridors. Across six well-known landmark structures in global cities, the image-based method achieves an overall detection accuracy of 87%, with a precision score of 68% for landmark-visible locations. In a second case study along the River Thames in London, the visibility graph reveals multi-landmark connections and identifies key mediating locations, with bridges accounting for approximately 31% of all connections. The proposed method complements LoS-based visibility analysis and offers a practical alternative in data-constrained settings. It also showcases the possibility of revealing the prevalent connections of visual objects in the urban environment, opening new perspectives for urban planning and heritage conservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。