arXiv:2506.23219cs.CVcs.AI2025-06ICCV被引 25

UrbanLLaVA让城市智能理解多模态数据,支持空间推理与跨场景应用。

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding

  • 构建跨模态城市指令数据集,融合多视角城市数据。
  • 多阶段训练提升空间推理能力,城市任务表现优于主流模型。
  • 适用于城市规划、交通分析等研究者,支持跨城市泛化。

城市研究涉及多种场景与任务,需理解多模态数据。现有方法多聚焦单一数据类型,缺乏统一框架。本文提出UrbanLLaVA,一种可同时处理四类城市数据的多模态大语言模型,在多样城市任务中表现优于通用MLLMs。我们构建了涵盖单模态与跨模态数据的城市指令数据集,覆盖从局部到全局的城市视图。提出分阶段训练框架,解耦空间推理增强与领域知识学习,提升模型兼容性与下游性能。同时扩展城市研究基准,评估模型在多种任务中的表现。三个城市的实验表明,UrbanLLaVA在单模态与复杂跨模态任务中均优于开源与专有MLLMs,具备强泛化能力。源码与数据已公开于https://github.com/tsinghua-fib-lab/UrbanLLaVA。

原文摘要 · Abstract (English)

Urban research involves a wide range of scenarios and tasks that require the understanding of multi-modal data. Current methods often focus on specific data types and lack a unified framework in urban field for processing them comprehensively. The recent success of multi-modal large language models (MLLMs) presents a promising opportunity to overcome this limitation. In this paper, we introduce $\textit{UrbanLLaVA}$, a multi-modal large language model designed to process these four types of data simultaneously and achieve strong performance across diverse urban tasks compared with general MLLMs. In $\textit{UrbanLLaVA}$, we first curate a diverse urban instruction dataset encompassing both single-modal and cross-modal urban data, spanning from location view to global view of urban environment. Additionally, we propose a multi-stage training framework that decouples spatial reasoning enhancement from domain knowledge learning, thereby improving the compatibility and downstream performance of $\textit{UrbanLLaVA}$ across diverse urban tasks. Finally, we also extend existing benchmark for urban research to assess the performance of MLLMs across a wide range of urban tasks. Experimental results from three cities demonstrate that $\textit{UrbanLLaVA}$ outperforms open-source and proprietary MLLMs in both single-modal tasks and complex cross-modal tasks and shows robust generalization abilities across cities. Source codes and data are openly accessible to the research community via https://github.com/tsinghua-fib-lab/UrbanLLaVA.

城市智能多模态空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。