arXiv:2505.18178cs.LGcs.AI2025-05

通过成对跨视图学习,让多模态区域表示更精准

Less is More: Multimodal Region Representation via Pairwise Inter-view Learning

  • 用成对交互学习分解多模态数据的共性和特异性信息
  • 在纽约和德里任务中表现优于现有方法,参数和计算量更低
  • 适合需要高效融合多源地理数据的研究者

随着地理空间数据的增多,区域表征学习(RRL)被用于分析复杂区域特征。现有方法多采用对比学习捕捉双模态间的共享信息,却忽略了每种模态特有的任务相关细节。这些特异性信息能解释仅靠共享信息无法捕捉的区域特性。引入信息分解可解决此问题,将多模态数据分解为共享与独特信息。但现有分解方法仅限于双模态,而RRL可利用多种地理数据。扩展至多模态面临高阶关系建模难题,导致学习目标呈组合爆炸,模型复杂度激增。本文提出交叉模态知识注入嵌入(CooKIE),一种面向RRL的信息分解方法,可同时捕捉共享与独特表征。CooKIE采用成对跨视图学习机制,在不建模高阶依赖的前提下捕获高阶信息,避免了穷举组合。我们在纽约市和印度德里开展三项回归任务和一项土地利用分类任务评估CooKIE。结果表明,该方法优于现有RRL方法及因子化模型,在减少训练参数与每秒浮点运算数(FLOPs)的同时,更有效捕捉多模态信息。代码已开源:https://github.com/MinNamgung/CooKIE。

原文摘要 · Abstract (English)

With the increasing availability of geospatial datasets, researchers have explored region representation learning (RRL) to analyze complex region characteristics. Recent RRL methods use contrastive learning (CL) to capture shared information between two modalities but often overlook task-relevant unique information specific to each modality. Such modality-specific details can explain region characteristics that shared information alone cannot capture. Bringing information factorization to RRL can address this by factorizing multimodal data into shared and unique information. However, existing factorization approaches focus on two modalities, whereas RRL can benefit from various geospatial data. Extending factorization beyond two modalities is non-trivial because modeling high-order relationships introduces a combinatorial number of learning objectives, increasing model complexity. We introduce Cross modal Knowledge Injected Embedding, an information factorization approach for RRL that captures both shared and unique representations. CooKIE uses a pairwise inter-view learning approach that captures high-order information without modeling high-order dependency, avoiding exhaustive combinations. We evaluate CooKIE on three regression tasks and a land use classification task in New York City and Delhi, India. Results show that CooKIE outperforms existing RRL methods and a factorized RRL model, capturing multimodal information with fewer training parameters and floating-point operations per second (FLOPs). We release the code: https://github.com/MinNamgung/CooKIE.

区域表征多模态学习信息分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。