UrbanFusion通过随机多模态融合,提升城市空间表征的鲁棒性与泛化能力。
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
- 采用分模态编码器+Transformer融合模块,实现多源地理数据统一表征。
- 在56个城市41项任务中超越主流模型,支持推理时灵活使用任意模态组合。
- 适用于数据不全场景,对未见区域也有良好适应性,适合城市智能应用开发。
预测房价、公共卫生等城市现象需有效融合多种地理空间数据。现有方法多依赖特定任务模型,而通用空间表征模型通常仅支持有限模态且缺乏多模态融合能力。为此,我们提出UrbanFusion,一种基于随机多模态融合(SMF)的空间表征模型。该框架使用分模态编码器处理街景图像、遥感数据、地图和兴趣点(POIs)等输入,并通过Transformer融合模块学习统一表征。在56个城市的41项任务上进行的广泛评估表明,UrbanFusion在性能和泛化能力上均优于当前最优的GeoAI模型:1)在位置编码任务中表现更优;2)推理时可灵活使用任意模态组合;3)对训练中未见区域具有良好的泛化能力。该模型可在预训练与推理阶段根据可用数据灵活选择任意模态子集,适用于多样化的数据可用性场景。
原文摘要 · Abstract (English)
Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent generic models for spatial representations often support only limited modalities and lack multimodal fusion capabilities. To overcome these challenges, we present UrbanFusion, a spatial representation model that features Stochastic Multimodal Fusion (SMF). The framework employs modality-specific encoders to process different types of inputs, including street view imagery, remote sensing data, cartographic maps, and points of interest (POIs) data. These multimodal inputs are integrated via a Transformer-based fusion module that learns unified representations. An extensive evaluation across 41 tasks in 56 cities worldwide demonstrates UrbanFusion's strong generalization and predictive performance compared to state-of-the-art GeoAI models. Specifically, it 1) outperforms prior models on location-encoding, 2) allows multimodal input during inference, and 3) generalizes well to regions unseen during training. UrbanFusion can flexibly utilize any subset of available modalities for a given location during both pretraining and inference, enabling broad applicability across diverse data availability scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。