arXiv:2512.22605cs.AIcs.CV2025-12被引 2

融合多模态时空知识,提升位置推荐的泛化能力。

Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation

  • 构建统一时空关系图,利用大模型增强的时空知识图谱表征多模态信息。
  • 设计门控机制融合多模态表示,通过STKG引导实现跨模态动态对齐。
  • 在6个公开数据集上表现优异,异常场景下仍具强泛化性,适合实际应用。

精准预测人类移动行为已产生显著社会经济效益,如位置推荐与疏散建议。然而现有方法泛化能力有限:单模态方法受限于数据稀疏性和固有偏差,多模态方法难以有效捕捉由静态多模态表示与时空动态间语义鸿沟导致的移动规律。为此,我们引入多模态时空知识以刻画移动动态,提出 extbf{M}ulti- extbf{M}odal extbf{Mob}ility ( extbf{M}$^3$ extbf{ob})。首先,基于大语言模型增强的时空知识图谱(STKG),构建统一的时空关系图(STRG)以表征多模态信息。其次,设计门控机制融合不同模态的时空图表示,并提出STKG引导的跨模态对齐,将时空动态知识注入静态图像模态。在六个公开数据集上的大量实验表明,所提方法不仅在正常场景中持续提升性能,且在异常场景下展现出显著泛化能力。

原文摘要 · Abstract (English)

The precise prediction of human mobility has produced significant socioeconomic impacts, such as location recommendations and evacuation suggestions. However, existing methods suffer from limited generalization capability: unimodal approaches are constrained by data sparsity and inherent biases, while multi-modal methods struggle to effectively capture mobility dynamics caused by the semantic gap between static multi-modal representation and spatial-temporal dynamics. Therefore, we leverage multi-modal spatial-temporal knowledge to characterize mobility dynamics for the location recommendation task, dubbed as \textbf{M}ulti-\textbf{M}odal \textbf{Mob}ility (\textbf{M}$^3$\textbf{ob}). First, we construct a unified spatial-temporal relational graph (STRG) for multi-modal representation, by leveraging the functional semantics and spatial-temporal knowledge captured by the large language models (LLMs)-enhanced spatial-temporal knowledge graph (STKG). Second, we design a gating mechanism to fuse spatial-temporal graph representations of different modalities, and propose an STKG-guided cross-modal alignment to inject spatial-temporal dynamic knowledge into the static image modality. Extensive experiments on six public datasets show that our proposed method not only achieves consistent improvements in normal scenarios but also exhibits significant generalization ability in abnormal scenarios.

位置推荐多模态时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。