构建全球城市时空基础模型,实现跨城市零样本泛化。
UrbanFM: Scaling Urban Spatio-Temporal Foundation Models

- 用数据、计算、架构三方面规模化构建城市时空模型
- 在超100个城市数据上训练,实现跨城市零样本预测
- 适合城市规划、交通管理等需要通用模型的场景
城市作为动态复杂系统,持续生成蕴含人类移动与城市发展规律的时空数据流。尽管科学人工智能在基因组学和气象学等领域已展现基础模型的变革力量,城市计算仍因‘场景特定’模型而碎片化,这些模型过度拟合于特定区域或任务,缺乏泛化能力。为弥合这一差距并推动城市时空基础模型发展,我们以规模化为核心视角,系统探究两个关键问题:规模化什么、如何规模化。基于第一性原理分析,我们识别出三个关键维度:异质性、相关性和动态性,并将其与城市时空数据的基本科学属性对齐。具体而言,为解决异质性,我们构建了涵盖超过100个全球城市的万亿级数据集WorldST,将交通流量、速度等多元物理信号统一为标准化格式;为实现计算规模化,提出MiniST单元,一种新型分块机制,将连续时空场离散为可学习计算单元,统一网格与传感器观测表示;为应对动态性,提出UrbanFM,一种极简自注意力架构,通过最小归纳偏置自主学习大规模数据中的动态时空依赖。此外,我们建立了目前最大规模的城市时空基准EvalST。大量实验表明,UrbanFM在未见过的城市和任务上实现显著零样本泛化,标志着迈向大规模城市时空基础模型的关键一步。
原文摘要 · Abstract (English)
Urban systems, as dynamic complex systems, continuously generate spatio-temporal data streams that encode the fundamental laws of human mobility and city evolution. While AI for Science has witnessed the transformative power of foundation models in disciplines like genomics and meteorology, urban computing remains fragmented due to "scenario-specific" models, which are overfitted to specific regions or tasks, hindering their generalizability. To bridge this gap and advance spatio-temporal foundation models for urban systems, we adopt scaling as the central perspective and systematically investigate two key questions: what to scale and how to scale. Grounded in first-principles analysis, we identify three critical dimensions: heterogeneity, correlation, and dynamics, aligning these principles with the fundamental scientific properties of urban spatio-temporal data. Specifically, to address heterogeneity through data scaling, we construct WorldST. This billion-scale corpus standardizes diverse physical signals, such as traffic flow and speed, from over 100 global cities into a unified data format. To enable computation scaling for modeling correlations, we introduce the MiniST unit, a novel split mechanism that discretizes continuous spatio-temporal fields into learnable computational units to unify representations of grid-based and sensor-based observations. Finally, addressing dynamics via architecture scaling, we propose UrbanFM, a minimalist self-attention architecture designed with limited inductive biases to autonomously learn dynamic spatio-temporal dependencies from massive data. Furthermore, we establish EvalST, the largest-scale urban spatio-temporal benchmark to date. Extensive experiments demonstrate that UrbanFM achieves remarkable zero-shot generalization across unseen cities and tasks, marking a pivotal first step toward large-scale urban spatio-temporal foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。