arXiv:2607.08357cs.AI2026-07被引 1

用离散扩散模型直接生成带语义的移动数据,更快更真实。

MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation

论文配图:MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation
图 1 · 摘自论文原文
  • 直接对多通道语义骨架去噪,跳过复杂中间步骤。
  • 在三地真实数据上保持轨迹长度和时间间隔分布,生成速度比现有方法快5.3倍。
  • 适合需要高效隐私保护数据生成的研究者和城市规划应用。

人类移动数据对交通优化、城市规划和资源分配至关重要,但真实数据采集成本高且因隐私问题难以共享。现有基于扩散的方法通常依赖连续或潜在时空轨迹,难以原生建模具有明确区域、活动、时间和区间结构的离散语义事件。为此,我们提出MobiDiff,一种端到端的离散扩散框架,通过直接对多通道语义骨架去噪来高效生成移动数据,避免了昂贵的插值、潜在轨迹构建和粗到细生成流程。具体而言,MobiDiff将每个出行事件分解为空间、活动和时间通道,并采用结构化事件、组和通道级掩码,联合捕捉轨迹级移动模式与事件内依赖关系。我们在亚特兰大、波士顿和西雅图三个大规模真实数据集上评估了生成保真度、隐私保护性和效率。结果表明,MobiDiff有效保留了轨迹长度和时间间隔分布,同时在更广泛的移动统计指标上保持竞争力;推理速度显著优于现有方法,平均比GeoGen快5.3倍。这些发现表明,离散扩散为合成移动数据提供了可解释且高效的框架。

原文摘要 · Abstract (English)

Human mobility data are essential for transportation optimization, urban planning, and resource allocation, yet real-world mobility data are costly to collect and difficult to share due to privacy concerns. Recent diffusion-based methods have shown promise in synthesizing realistic mobility patterns, but they typically rely on continuous or latent spatio-temporal traces, limiting their ability to natively model discrete semantic events with explicit region, activity, time, and interval structures. To address this issue, we introduce MobiDiff, an end-to-end discrete diffusion framework that efficiently generates mobility data by directly denoising multi-channel semantic skeletons, avoiding the costly interpolation, latent trace construction, and coarse-to-fine realization pipelines widely used in existing diffusion-based methods. Specifically, MobiDiff decomposes each human check-in event into spatial, activity, and temporal channels, and employs structured event-, group-, and channel-level masking to jointly capture trajectory-level mobility patterns and within-event dependencies. We evaluate generation fidelity, privacy-preserving, and efficiency on three large-scale real-world datasets from Atlanta, Boston, and Seattle. Results show that MobiDiff effectively preserves trajectory length and temporal interval distributions while remaining competitive across broader mobility statistics; it is also much faster than state-of-the-art methods, e.g., 5.3$\times$ faster than GeoGen on average during inference. These findings suggest that discrete diffusion offers an interpretable and efficient framework for synthetic mobility data generation.

移动数据生成扩散模型隐私保护离散扩散

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。