arXiv:2509.23772cs.CVstat.AP2025-09

针对城市区域建模中多模态数据融合不精准的问题,提出自适应分层图模型。

A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning

  • 按空间密度将多源数据分为聚合与点级两类,分别用专家网络和双层图网络处理
  • 动态学习不同区域的模态融合权重,提升复杂空间下的表征能力
  • 在六种模态、三个任务上超越现有方法,适合城市计算与智能规划研究者

基于图的模型已成为建模多模态城市数据并学习区域表示的有效范式。然而,现有方法存在两大局限:(1) 通常对所有模态使用相同的图神经网络架构,无法捕捉模态特异性结构;(2) 融合阶段常忽略空间异质性,假设模态聚合权重在各区域间不变,导致表征效果不佳。为此,本文提出MTGRR框架,基于包含兴趣点(POI)、出租车轨迹、土地利用、道路要素、遥感影像和街景图像的多模态数据集。首先,根据空间密度与数据特性将模态分为聚合级与点级两类:聚合级采用混合专家(MoE)图架构,每类模态由专用专家GNN处理以捕获其特异性特征;点级则构建双层图网络提取细粒度视觉语义。其次,设计空间感知的多模态融合机制,动态推断区域特定的模态融合权重。在此基础上,引入联合对比学习策略,整合聚合级、点级及融合级目标,优化区域表示。在两个真实世界数据集上的实验表明,该框架在六种模态、三种任务中均持续优于现有最先进方法,验证了其有效性。

原文摘要 · Abstract (English)

Graph-based models have emerged as a powerful paradigm for modeling multimodal urban data and learning region representations for various downstream tasks. However, existing approaches face two major limitations. (1) They typically employ identical graph neural network architectures across all modalities, failing to capture modality-specific structures and characteristics. (2) During the fusion stage, they often neglect spatial heterogeneity by assuming that the aggregation weights of different modalities remain invariant across regions, resulting in suboptimal representations. To address these issues, we propose MTGRR, a modality-tailored graph modeling framework for urban region representation, built upon a multimodal dataset comprising point of interest (POI), taxi mobility, land use, road element, remote sensing, and street view images. (1) MTGRR categorizes modalities into two groups based on spatial density and data characteristics: aggregated-level and point-level modalities. For aggregated-level modalities, MTGRR employs a mixture-of-experts (MoE) graph architecture, where each modality is processed by a dedicated expert GNN to capture distinct modality-specific characteristics. For the point-level modality, a dual-level GNN is constructed to extract fine-grained visual semantic features. (2) To obtain effective region representations under spatial heterogeneity, a spatially-aware multimodal fusion mechanism is designed to dynamically infer region-specific modality fusion weights. Building on this graph modeling framework, MTGRR further employs a joint contrastive learning strategy that integrates region aggregated-level, point-level, and fusion-level objectives to optimize region representations. Experiments on two real-world datasets across six modalities and three tasks demonstrate that MTGRR consistently outperforms state-of-the-art baselines, validating its effectiveness.

城市计算图神经网络多模态融合对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。