arXiv:2510.13774cs.LGcs.CV2025-10被引 3

UrbanFusion通过随机多模态融合,提升城市空间表征的鲁棒性与泛化能力。

UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations

  • 采用分模态编码器+Transformer融合模块,实现多源地理数据统一表征。
  • 在56个城市41项任务中超越主流模型,支持推理时灵活使用任意模态组合。
  • 适用于数据不全场景,对未见区域也有良好适应性,适合城市智能应用开发。

预测房价、公共卫生等城市现象需有效融合多种地理空间数据。现有方法多依赖特定任务模型,而通用空间表征模型通常仅支持有限模态且缺乏多模态融合能力。为此,我们提出UrbanFusion,一种基于随机多模态融合(SMF)的空间表征模型。该框架使用分模态编码器处理街景图像、遥感数据、地图和兴趣点(POIs)等输入,并通过Transformer融合模块学习统一表征。在56个城市的41项任务上进行的广泛评估表明,UrbanFusion在性能和泛化能力上均优于当前最优的GeoAI模型:1)在位置编码任务中表现更优;2)推理时可灵活使用任意模态组合;3)对训练中未见区域具有良好的泛化能力。该模型可在预训练与推理阶段根据可用数据灵活选择任意模态子集,适用于多样化的数据可用性场景。

原文摘要 · Abstract (English)

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent generic models for spatial representations often support only limited modalities and lack multimodal fusion capabilities. To overcome these challenges, we present UrbanFusion, a spatial representation model that features Stochastic Multimodal Fusion (SMF). The framework employs modality-specific encoders to process different types of inputs, including street view imagery, remote sensing data, cartographic maps, and points of interest (POIs) data. These multimodal inputs are integrated via a Transformer-based fusion module that learns unified representations. An extensive evaluation across 41 tasks in 56 cities worldwide demonstrates UrbanFusion's strong generalization and predictive performance compared to state-of-the-art GeoAI models. Specifically, it 1) outperforms prior models on location-encoding, 2) allows multimodal input during inference, and 3) generalizes well to regions unseen during training. UrbanFusion can flexibly utilize any subset of available modalities for a given location during both pretraining and inference, enabling broad applicability across diverse data availability scenarios.

空间表征多模态融合城市智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。