构建统一城市多模态嵌入模型,提升复杂地理任务理解能力
Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

- 用统一框架处理街景、遥感、文本等异构地理数据
- 在45项城市任务上比最强基线提升15.3%准确率
- 适合做城市分析、智能导航与时空变化检测的研究者
地理空间与城市应用日益需要模型在街景图像、遥感观测、文本描述、区域建议和时序变化线索之间进行异构信息对比。然而现有多模态嵌入模型与评测基准仍主要围绕通用图像-文本匹配设计,未明确统一嵌入空间是否能支持涉及空间关系、细粒度语义和时序变化的复杂地理任务。为此,我们提出三项关键贡献:第一,构建GeoMEB——一个大规模多模态嵌入基准,涵盖45个城市评估任务(包括检索、视觉问答、变化检测、分类与视觉定位),包含132万训练样本和28.6万查询;第二,提出Geo-Embed,一种统一嵌入模型,通过共享视觉-语言主干网络,实现对单图、多图、文本、区域与掩码等异构地理输入的指令条件化查询-目标匹配。在GeoMEB上,Geo-Embed在代表性多模态嵌入器中表现最优,相较最强基线相对提升15.3%。结果表明未来地理嵌入模型应围绕显式的查询-目标关系(语义、跨视图、区域级、时序对应)组织训练与评测。
原文摘要 · Abstract (English)
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。