arXiv:2605.06990cs.CVcs.LG2026-05

让轨迹与街景图像精细对齐,提升城市移动理解能力

TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

论文配图:TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
图 1 · 摘自论文原文
  • 用连续神经表示建模轨迹,实现任意点与街景对齐
  • 在4个城市任务中超越现有方法,显著提升轨迹理解性能
  • 适合研究城市交通、导航或时空多模态学习的学者

多模态自监督学习(MSSL)已成为预训练地理空间基础模型的关键范式。然而,现有地理空间MSSL方法主要针对静态模态对(如卫星影像、街景图像和文本),通过对齐同一或邻近位置的观测来学习。这一假设在人类移动轨迹上失效——轨迹是沿路径的连续运动,而非单点观测。尽管轨迹能反映随时间在道路、街区和场所间的人类活动,却在当前地理空间MSSL框架中仍被忽视。我们提出TrajGANR,一种以轨迹为中心的地理空间多模态自监督框架,将连续移动模式与静态位置观测对齐。TrajGANR在路径任意点学习轨迹的连续神经表示,实现与附近街景图像的细粒度对齐,即使这些街景不位于轨迹关键点。我们据此设计了一种联合对齐三种模态(轨迹、街景图像、地理位置)的MSSL目标。在四个城市移动与道路理解任务上评估,TrajGANR始终优于现有地理空间MSSL框架及专门的轨迹基础模型。消融实验表明,所提MSSL目标与多模态学习框架是性能提升的主要驱动力,凸显细粒度地理对齐与多模态学习的重要性。

原文摘要 · Abstract (English)

Multimodal self-supervised learning (MSSL) has emerged as a key paradigm for pretraining geospatial foundation models. However, existing geospatial MSSL methods are mainly designed for static pairs of modalities, such as satellite imagery, street-view imagery, and text, where learning is driven by aligning observations from the same or nearby locations. This assumption breaks down for human mobility trajectories, which represent continuous movement along paths rather than discrete observations at individual locations. Although trajectories are important for urban understanding through their ability to capture human activity across roads, neighborhoods, and places over time, they remain largely underexplored in current geospatial MSSL frameworks. We present TrajGANR, a novel trajectory-centric geospatial MSSL framework that aligns continuous movement patterns with static, location-based observations. TrajGANR learns a continuous neural representation of trajectories at arbitrary points along each path, which enables fine-grained alignment with nearby street-view images, even when they are not co-located with any trajectory waypoints. We leverage this capability to introduce an MSSL objective that jointly aligns three modalities: trajectories, street-view images, and their geographic locations. We evaluate TrajGANR on four urban mobility and road understanding tasks. Across these tasks, TrajGANR consistently outperforms existing geospatial MSSL frameworks and a trajectory-specific foundation model. Ablation studies further demonstrate that our proposed MSSL objective and the multimodal learning framework are the primary drivers of these improvements, highlighting the importance of fine-grained geospatial alignment over coarser aggregation, as well as geospatial multimodal learning.

轨迹学习多模态地理空间自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。