arXiv:2511.16105cs.LG2025-11

用路径片段字典学习生成高效鲁棒的轨迹,适合噪声数据和隐私保护场景。

Data-Efficient and Robust Trajectory Generation through Pathlet Dictionary Learning

  • 基于路径片段字典的二值编码,结合变分自编码器建模轨迹生成过程。
  • 在真实数据集上比基线提升35.4%和26.3%,且能处理噪声数据。
  • 生成轨迹可直接用于预测与去噪,训练效率提升64.8%和56.5%。

轨迹生成在隐私保护的城市出行研究与位置服务中日益受到关注。尽管已有大量研究采用深度学习或生成式AI方法建模轨迹并取得良好效果,但这些模型的鲁棒性与可解释性仍缺乏探索,限制了其在嘈杂真实数据上的应用及下游任务中的可信度。为此,本文利用城市轨迹中的规律结构,提出一种基于路径片段(pathlet)表示的深度生成模型,通过二值向量与学习得到的片段字典对轨迹进行编码。具体而言,设计了一种概率图模型,包含变分自编码器(VAE)与线性解码器组件。训练时可同步学习路径片段表示的隐变量与捕捉轨迹数据集中的移动模式的字典。模型的条件版本可基于时空约束生成定制化轨迹。实验表明,该模型即使在噪声数据下也能有效学习数据分布,在两个真实轨迹数据集上分别相对基线提升35.4%和26.3%。生成的轨迹可直接用于轨迹预测与数据去噪等下游任务。此外,框架设计带来显著效率优势,相比先前方法节省64.8%时间与56.5% GPU内存。

原文摘要 · Abstract (English)

Trajectory generation has recently drawn growing interest in privacy-preserving urban mobility studies and location-based service applications. Although many studies have used deep learning or generative AI methods to model trajectories and have achieved promising results, the robustness and interpretability of such models are largely unexplored. This limits the application of trajectory generation algorithms on noisy real-world data and their trustworthiness in downstream tasks. To address this issue, we exploit the regular structure in urban trajectories and propose a deep generative model based on the pathlet representation, which encode trajectories with binary vectors associated with a learned dictionary of trajectory segments. Specifically, we introduce a probabilistic graphical model to describe the trajectory generation process, which includes a Variational Autoencoder (VAE) component and a linear decoder component. During training, the model can simultaneously learn the latent embedding of pathlet representations and the pathlet dictionary that captures mobility patterns in the trajectory dataset. The conditional version of our model can also be used to generate customized trajectories based on temporal and spatial constraints. Our model can effectively learn data distribution even using noisy data, achieving relative improvements of $35.4\%$ and $26.3\%$ over strong baselines on two real-world trajectory datasets. Moreover, the generated trajectories can be conveniently utilized for multiple downstream tasks, including trajectory prediction and data denoising. Lastly, the framework design offers a significant efficiency advantage, saving $64.8\%$ of the time and $56.5\%$ of GPU memory compared to previous approaches.

轨迹生成路径片段生成模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。