arXiv:2606.15890cs.AI2026-06KDD

构建首个城市福祉多模态推理基准,评估模型对时空城市状态的理解能力。

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

论文配图:UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
图 1 · 摘自论文原文
  • 基于卫星与街景图像联合建模,统一网格化评估城市多维度福祉指标
  • 覆盖38城多年数据,涵盖环境、可达性、城市形态等5类指标,支持时序预测与趋势分类
  • 首次系统评测15个前沿多模态大模型在真实城市场景中的时空推理表现

从多模态数据理解城市福祉需整合异构的空间与时间信号,对当前多模态大语言模型(MLLMs)构成重大挑战。本文提出UrbanWell,一个大规模基准,通过联合建模卫星图像与街景图像,系统评估MLLMs在城市福祉分析中的时空推理能力。UrbanWell覆盖全球38个城市,跨多个年份,包含五类指标:(1) 环境条件(CO₂、NO₂、PM₂.₅ 和归一化植被指数),(2) 空间可达性(至超市和餐厅的最短距离),(3) 城市形态(道路长度、道路密度、土地利用),(4) 城市活力(人口、经济活动多样性、土地利用多样性),(5) 主观感知属性(如安全、美丽、活力、富裕、安静)。所有指标均在网格层面对齐,实现标准化评估。除静态预测外,还定义了未来值预测与时间趋势分类等时序推理任务。我们在零样本设置下对15个代表性先进MLLMs进行评测,结果表明,尽管模型能捕捉显著的空间与感知线索,但在环境与主观感知等异构城市指标上的表现差异显著。UrbanWell为城市福祉分析中的多模态空间与时间推理提供统一基准,构建了系统评估与未来研究的标准化测试平台。代码与数据集可通过 https://github.com/axin1301/UrbanWell-Benchmark 获取。

原文摘要 · Abstract (English)

Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWell, a large-scale benchmark designed to systematically evaluate the spatio-temporal reasoning capabilities of MLLMs for urban wellbeing analytics through joint modeling of satellite and street view imagery. UrbanWell spans 38 cities across multiple years and includes diverse indicators covering (1) environmental conditions (CO$_2$, NO$_2$, PM${2.5}$, and Normalized Difference Vegetation Index), (2) spatial accessibility (minimum distance to supermarkets and restaurants), (3) urban form (road length, road density, and land use), (4) urban vitality (population, economic activity diversity, and land use diversity), and (5) subjective perception attributes (e.g., safety, beauty, liveliness, wealth, and quietness). All indicators are aligned at grid level to enable standardized evaluation. Beyond static prediction, UrbanWell defines temporal reasoning tasks, including future value forecasting from historical observations and temporal trend classification. We benchmark 15 state-of-the-art representative MLLMs in a zero-shot setting, providing a comprehensive comparative evaluation across spatial and temporal dimensions. Experimental results indicate that while MLLMs capture salient spatial and perceptual cues, their performance varies substantially across heterogeneous urban indicators spanning environment and subjective perception. UrbanWell serves as a unified benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics, offering a standardized testbed for systematic assessment and future research on multimodal urban intelligence. Our codes and datasets are accessible via https://github.com/axin1301/UrbanWell-Benchmark.

城市福祉多模态时空推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。