arXiv:2502.04638cs.CVcs.AI2025-02被引 5

通过时空对比学习,让街景图像更好理解城市动态与环境氛围。

Learning Street View Representations with Spatiotemporal Contrast

  • 利用同一位置不同时段、相邻视角的街景图构建对比任务。
  • 在视觉定位、社会经济估算等任务上显著超越传统方法。
  • 适合城市规划、环境感知等领域的研究者参考使用。

街景图像广泛应用于城市视觉环境的表征学习,支持环境感知与社会经济评估等可持续发展任务。然而,现有图像表征难以有效捕捉街景中动态环境(如行人、车辆、植被)、建成环境(如建筑、道路、基础设施)及环境氛围(如文化、社会经济气息)。本文提出一种创新的自监督学习框架,利用街景图像的时空特性,学习动态城市环境的表征以支持多样下游任务。通过在同一位置不同时段采集的图像以及同一时刻空间邻近视角的图像,构建对比学习任务,旨在学习建成环境的时不变特征和邻里氛围的空间不变特征。实验表明,该方法在视觉地点识别、社会经济估计和人-环境感知等任务上显著优于传统有监督与无监督方法。此外,我们展示了不同对比学习目标所学表征在各类下游任务中的差异表现。本研究系统探讨了基于街景图像的城市表征学习策略,为视觉数据在城市科学中的应用提供了基准。代码已开源:https://github.com/yonglleee/UrbanSTCL。

原文摘要 · Abstract (English)

Street view imagery is extensively utilized in representation learning for urban visual environments, supporting various sustainable development tasks such as environmental perception and socio-economic assessment. However, it is challenging for existing image representations to specifically encode the dynamic urban environment (such as pedestrians, vehicles, and vegetation), the built environment (including buildings, roads, and urban infrastructure), and the environmental ambiance (such as the cultural and socioeconomic atmosphere) depicted in street view imagery to address downstream tasks related to the city. In this work, we propose an innovative self-supervised learning framework that leverages temporal and spatial attributes of street view imagery to learn image representations of the dynamic urban environment for diverse downstream tasks. By employing street view images captured at the same location over time and spatially nearby views at the same time, we construct contrastive learning tasks designed to learn the temporal-invariant characteristics of the built environment and the spatial-invariant neighborhood ambiance. Our approach significantly outperforms traditional supervised and unsupervised methods in tasks such as visual place recognition, socioeconomic estimation, and human-environment perception. Moreover, we demonstrate the varying behaviors of image representations learned through different contrastive learning objectives across various downstream tasks. This study systematically discusses representation learning strategies for urban studies based on street view images, providing a benchmark that enhances the applicability of visual data in urban science. The code is available at https://github.com/yonglleee/UrbanSTCL.

街景表征对比学习城市科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。