arXiv:2602.08342cs.CVcs.AI2026-02被引 1

为城市科学构建可迁移的空间对齐多模态嵌入,提升图像与地理结构的关联理解。

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

  • 基于空间图结构对齐街景图像与城市结构,引入路径推理和空间描述作为监督信号。
  • 在训练城市上图像检索提升44%,地理定位排名提升30%,跨城市泛化性能显著。
  • 适合研究城市感知、地理定位与空间理解的多模态模型开发者。

学习可迁移的城市多模态嵌入面临挑战,因城市理解具有本质空间性,而现有数据集缺乏街景图像与城市结构的显式对齐。我们提出UGData,一个空间对齐的数据集,将街景图像锚定在结构化空间图上,并通过空间推理路径和空间上下文描述提供图对齐的监督信号,揭示距离、方向、连通性及邻里上下文信息。在此基础上,我们提出UGE,一种两阶段训练策略,结合指令引导对比学习与基于图的空间编码,逐步稳定地对齐图像、文本与空间结构。我们进一步构建UGBench,用于评估空间嵌入在多种城市理解任务中的表现,包括地理定位排序、图像检索、城市感知和空间定位。我们在多个先进视觉语言模型(VLM)基线(Qwen2-VL、Qwen2.5-VL、Phi-3-Vision、LLaVA1.6-Mistral)上实现固定维度空间嵌入,采用LoRA微调。以Qwen2.5-VL-7B为基底的UGE,在训练城市上图像检索最高提升44%,地理定位排名提升30%;在未见城市上分别取得超过30%和22%的增益,验证了显式空间对齐在空间密集型城市任务中的有效性。

原文摘要 · Abstract (English)

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban structure. We introduce UGData, a spatially grounded dataset that anchors street-view images to structured spatial graphs and provides graph-aligned supervision via spatial reasoning paths and spatial context captions, exposing distance, directionality, connectivity, and neighborhood context beyond image content. Building on UGData, we propose UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding. We finally introduce UGBench, a comprehensive benchmark to evaluate how spatially grounded embeddings support diverse urban understanding tasks -- including geolocation ranking, image retrieval, urban perception, and spatial grounding. We develop UGE on multiple state-of-the-art VLM backbones, including Qwen2-VL, Qwen2.5-VL, Phi-3-Vision, and LLaVA1.6-Mistral, and train fixed-dimensional spatial embeddings with LoRA tuning. UGE built upon Qwen2.5-VL-7B backbone achieves up to 44% improvement in image retrieval and 30% in geolocation ranking on training cities, and over 30% and 22% gains respectively on held-out cities, demonstrating the effectiveness of explicit spatial grounding for spatially intensive urban tasks.

多模态空间建模城市科学视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。