用历史轨迹预测个人未来数日行程,提升隐私保护下的建模效果。
Training Machine Learning Models on Human Spatio-temporal Mobility Data: An Experimental Study [Experiment Paper]
- 引入日期、用户历史等语义信息增强模型对生活规律的理解。
- 小批量随机梯度优化在数据有限时显著提升预测精度。
- 采用分层采样与用户聚类缓解数据偏差,保障样本代表性。
个体级人类移动预测已成为传染病监测、老人与儿童照护等领域的重要研究方向。现有研究多聚焦于短期轨迹或下一位置预测,对宏观移动模式与生活规律关注不足。本文针对这一空白,系统实验分析多种模型、参数配置与训练策略,结合人类移动模式的统计分布特性,评估长短期记忆网络与基于Transformer的架构在预测个体未来数天至数周完整轨迹上的表现。结果表明,显式引入星期几、用户历史等语义信息可显著提升模型对个人生活规律的理解能力。由于隐私限制常导致用户信息缺失,我们发现用户采样不当会加剧数据偏斜,造成预测性能大幅下降。为此,提出基于用户语义聚类的分层采样方法以维持数据多样性与代表性。实验还表明,在数据量有限时,小批量随机梯度优化能有效提升模型性能。
原文摘要 · Abstract (English)
Individual-level human mobility prediction has emerged as a significant topic of research with applications in infectious disease monitoring, child, and elderly care. Existing studies predominantly focus on the microscopic aspects of human trajectories: such as predicting short-term trajectories or the next location visited, while offering limited attention to macro-level mobility patterns and the corresponding life routines. In this paper, we focus on an underexplored problem in human mobility prediction: determining the best practices to train a machine learning model using historical data to forecast an individuals complete trajectory over the next days and weeks. In this experiment paper, we undertake a comprehensive experimental analysis of diverse models, parameter configurations, and training strategies, accompanied by an in-depth examination of the statistical distribution inherent in human mobility patterns. Our empirical evaluations encompass both Long Short-Term Memory and Transformer-based architectures, and further investigate how incorporating individual life patterns can enhance the effectiveness of the prediction. We show that explicitly including semantic information such as day-of-the-week and user-specific historical information can help the model better understand individual patterns of life and improve predictions. Moreover, since the absence of explicit user information is often missing due to user privacy, we show that the sampling of users may exacerbate data skewness and result in a substantial loss in predictive accuracy. To mitigate data imbalance and preserve diversity, we apply user semantic clustering with stratified sampling to ensure that the sampled dataset remains representative. Our results further show that small-batch stochastic gradient optimization improves model performance, especially when human mobility training data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。