用语义化位置标记提升大模型对移动行为的分析能力
Enhancing Large Language Models for Mobility Analytics with Semantic Location Tokenization
- 将地理位置转化为语义丰富的小型标记,更好表达位置含义
- 在三个真实数据集上,预测下一步位置和恢复移动轨迹效果更优
- 适合做城市出行分析、轨迹建模的研究者与应用开发者
基于位置的服务广泛应用催生了海量移动数据,为建模城市中用户移动动态提供了机遇。当前方法主要存在两大局限:位置仅用离散编号表示,语义信息不足;且大语言模型(LLMs)对移动信号建模能力有限,依赖单一模板指令微调。为此,我们提出QT-Mob框架,通过引入位置标记模块,学习紧凑而富含语义的位置表示,保留上下文信息并兼容大模型。同时,设计多重互补微调目标,使学习到的标记与大模型内部表征对齐,增强对序列移动模式与位置语义的理解。实验在三个真实数据集上验证,该框架在下一步位置预测与移动轨迹恢复任务中均优于现有深度学习及基于大模型的方法。
原文摘要 · Abstract (English)
The widespread adoption of location-based services has led to the generation of vast amounts of mobility data, providing significant opportunities to model user movement dynamics within urban environments. Recent advancements have focused on adapting Large Language Models (LLMs) for mobility analytics. However, existing methods face two primary limitations: inadequate semantic representation of locations (i.e., discrete IDs) and insufficient modeling of mobility signals within LLMs (i.e., single templated instruction fine-tuning). To address these issues, we propose QT-Mob, a novel framework that significantly enhances LLMs for mobility analytics. QT-Mob introduces a location tokenization module that learns compact, semantically rich tokens to represent locations, preserving contextual information while ensuring compatibility with LLMs. Furthermore, QT-Mob incorporates a series of complementary fine-tuning objectives that align the learned tokens with the internal representations in LLMs, improving the model's comprehension of sequential movement patterns and location semantics. The proposed QT-Mob framework not only enhances LLMs' ability to interpret mobility data but also provides a more generalizable approach for various mobility analytics tasks. Experiments on three real-world dataset demonstrate the superior performance in both next-location prediction and mobility recovery tasks, outperforming existing deep learning and LLM-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。