arXiv:2506.02744cs.CEcs.AI2025-06被引 5

用兴趣点名称增强城市位置表征,提升分类与地图精度

Enriching Location Representation with Detailed Semantic Information

  • 将兴趣点名称与类别标签融合进多模态对比学习框架
  • 在用地分类和收入分布映射任务中提升4%至11%
  • 适合城市建模、地理信息分析的研究者使用

能同时捕捉城市环境结构与语义特征的空间表征对城市建模至关重要。传统空间嵌入通常侧重空间邻近性,而忽视了场所的细粒度上下文信息。为此,我们提出CaLLiPer+,作为CaLLiPer模型的扩展,通过多模态对比学习框架系统地整合兴趣点(POI)名称与类别标签。我们在两个下游任务——用地分类与社会经济地位分布制图——上评估其有效性,结果显示性能相比基线方法持续提升4%至11%。此外,引入POI名称显著提升了位置检索能力,使模型能更精准捕捉复杂城市概念。消融实验进一步揭示了POI名称与类别标签的互补作用,以及利用预训练文本编码器对空间表征的优势。总体而言,研究结果表明,融合细粒度语义属性与多模态学习技术,有助于推动城市基础模型的发展。

原文摘要 · Abstract (English)

Spatial representations that capture both structural and semantic characteristics of urban environments are essential for urban modeling. Traditional spatial embeddings often prioritize spatial proximity while underutilizing fine-grained contextual information from places. To address this limitation, we introduce CaLLiPer+, an extension of the CaLLiPer model that systematically integrates Point-of-Interest (POI) names alongside categorical labels within a multimodal contrastive learning framework. We evaluate its effectiveness on two downstream tasks, land use classification and socioeconomic status distribution mapping, demonstrating consistent performance gains of 4% to 11% over baseline methods. Additionally, we show that incorporating POI names enhances location retrieval, enabling models to capture complex urban concepts with greater precision. Ablation studies further reveal the complementary role of POI names and the advantages of leveraging pretrained text encoders for spatial representations. Overall, our findings highlight the potential of integrating fine-grained semantic attributes and multimodal learning techniques to advance the development of urban foundation models.

城市建模多模态学习位置表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。