arXiv:2508.03042cs.LG2025-08被引 2

用扩散模型统一城市画像的预训练与推理,提升未知区域预测精度

Urban In-Context Learning: Bridging Pretraining and Inference through Masked Diffusion for Urban Profiling

  • 通过掩码扩散机制实现城市数据的端到端建模
  • 在两个城市的三项指标上超越现有两阶段方法
  • 适合城市规划、智能交通等需精准区域预测的场景

城市画像旨在预测未知区域的城市特征,在经济和社会普查中具有关键作用。现有方法多采用两阶段范式:先学习城市表征,再通过线性探测进行下游预测,源于BERT时代。受GPT类模型启发,近年研究发现新型自监督预训练可使模型直接适用于下游任务,避免特定任务微调。这主要得益于GPT通过下一项预测统一预训练与推理形式。然而,城市数据在结构上与语言有本质差异,难以设计出统一预训练与推理的一阶段模型。本文提出城市上下文学习框架,通过城市区域上的掩码自编码过程统一预训练与推理。为捕捉城市画像分布,引入城市掩码扩散变换器,使每个区域的预测以概率分布形式表示而非确定值。此外,为稳定扩散训练,提出城市表征对齐机制,通过与经典城市画像方法的中间特征对齐来正则化模型。在两个城市三个指标上的大量实验表明,该一阶段方法持续优于最先进两阶段方法。消融实验与案例研究进一步验证了各模块有效性,尤其是扩散建模的作用。

原文摘要 · Abstract (English)

Urban profiling aims to predict urban profiles in unknown regions and plays a critical role in economic and social censuses. Existing approaches typically follow a two-stage paradigm: first, learning representations of urban areas; second, performing downstream prediction via linear probing, which originates from the BERT era. Inspired by the development of GPT style models, recent studies have shown that novel self-supervised pretraining schemes can endow models with direct applicability to downstream tasks, thereby eliminating the need for task-specific fine-tuning. This is largely because GPT unifies the form of pretraining and inference through next-token prediction. However, urban data exhibit structural characteristics that differ fundamentally from language, making it challenging to design a one-stage model that unifies both pretraining and inference. In this work, we propose Urban In-Context Learning, a framework that unifies pretraining and inference via a masked autoencoding process over urban regions. To capture the distribution of urban profiles, we introduce the Urban Masked Diffusion Transformer, which enables each region' s prediction to be represented as a distribution rather than a deterministic value. Furthermore, to stabilize diffusion training, we propose the Urban Representation Alignment Mechanism, which regularizes the model's intermediate features by aligning them with those from classical urban profiling methods. Extensive experiments on three indicators across two cities demonstrate that our one-stage method consistently outperforms state-of-the-art two-stage approaches. Ablation studies and case studies further validate the effectiveness of each proposed module, particularly the use of diffusion modeling.

城市画像扩散模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。