arXiv:2506.00530cs.AIcs.CL2025-06中稿 · ICLR被引 8

用卫星和街景图评估大模型预测城市经济状况的能力。

CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing

  • 构建17城多模态数据集,覆盖6大领域11项预测任务。
  • 17个顶尖大模型在跨域任务中表现参差,普遍难准确预测指标。
  • 为城市政策研究者提供评估大模型能力的统一框架。

通过视觉数据理解城市经济社会状况对可持续城市发展和政策规划至关重要。本文提出CityLens,一个综合性基准,用于评估大视觉语言模型(LVLMs)从卫星与街景图像中预测经济社会指标的能力。我们构建了覆盖全球17座城市的多模态数据集,涵盖经济、教育、犯罪、交通、健康与环境6个关键领域,反映城市生活的复杂性。基于此数据集,定义11项预测任务,并采用三种评估范式:直接指标预测、归一化指标估计与基于特征的回归。我们在这些任务上对17个前沿LVLM进行了基准测试。结果表明,尽管LVLM具备良好的感知与推理能力,但在预测城市经济社会指标方面仍存在局限。CityLens提供了诊断这些局限性的统一框架,有助于未来利用LVLM理解与预测城市经济社会模式。代码与数据已公开于https://github.com/tsinghua-fib-lab/CityLens。

原文摘要 · Abstract (English)

Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introduce \textit{CityLens}, a comprehensive benchmark designed to evaluate the capabilities of Large Vision-Language Models (LVLMs) in predicting socioeconomic indicators from satellite and street view imagery. We construct a multi-modal dataset covering a total of 17 globally distributed cities, spanning 6 key domains: economy, education, crime, transport, health, and environment, reflecting the multifaceted nature of urban life. Based on this dataset, we define 11 prediction tasks and utilize 3 evaluation paradigms: Direct Metric Prediction, Normalized Metric Estimation, and Feature-Based Regression. We benchmark 17 state-of-the-art LVLMs across these tasks. These make CityLens the most extensive socioeconomic benchmark to date in terms of geographic coverage, indicator diversity, and model scale. Our results reveal that while LVLMs demonstrate promising perceptual and reasoning capabilities, they still exhibit limitations in predicting urban socioeconomic indicators. CityLens provides a unified framework for diagnosing these limitations and guiding future efforts in using LVLMs to understand and predict urban socioeconomic patterns. The code and data are available at https://github.com/tsinghua-fib-lab/CityLens.

城市感知大模型评估多模态社会计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。