用大模型预测城市动态,兼顾准确与泛化能力
UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models
- 设计多模态时空掩码自编码器,捕捉城市数据复杂关联
- 在多城真实数据上实现比现有方法更高的预测精度
- 支持零样本预测,适合跨城市应用的智慧城市场景
理解与预测城市动态对交通管理、城市规划和公共服务优化至关重要。尽管神经网络方法已取得进展,但通常依赖特定任务架构和大量数据,难以跨不同城市场景泛化。大型语言模型虽具强推理与泛化能力,但在时空城市动态预测中的应用仍不足。现有方法难以有效整合多源时空数据,且无法应对训练与测试数据分布差异,影响实际预测可靠性。为此,我们提出UrbanMind,一种面向多维度城市动态预测的新型时空大模型框架,兼具高精度与鲁棒泛化能力。核心是Muffin-MAE——一种具备专用掩码策略的多模态融合掩码自编码器,可捕捉复杂时空依赖与多维城市动态间的相互关系。同时设计语义感知提示与微调策略,将时空上下文编码进提示中,增强大模型对时空模式的推理能力。为进一步提升泛化性,引入测试时自适应机制,通过测试数据重建器动态调整大模型生成的嵌入表示。在多个城市的实测数据集上广泛实验表明,UrbanMind持续优于现有最优基线,在零样本设置下仍保持高精度与强泛化能力。
原文摘要 · Abstract (English)
Understanding and predicting urban dynamics is crucial for managing transportation systems, optimizing urban planning, and enhancing public services. While neural network-based approaches have achieved success, they often rely on task-specific architectures and large volumes of data, limiting their ability to generalize across diverse urban scenarios. Meanwhile, Large Language Models (LLMs) offer strong reasoning and generalization capabilities, yet their application to spatial-temporal urban dynamics remains underexplored. Existing LLM-based methods struggle to effectively integrate multifaceted spatial-temporal data and fail to address distributional shifts between training and testing data, limiting their predictive reliability in real-world applications. To bridge this gap, we propose UrbanMind, a novel spatial-temporal LLM framework for multifaceted urban dynamics prediction that ensures both accurate forecasting and robust generalization. At its core, UrbanMind introduces Muffin-MAE, a multifaceted fusion masked autoencoder with specialized masking strategies that capture intricate spatial-temporal dependencies and intercorrelations among multifaceted urban dynamics. Additionally, we design a semantic-aware prompting and fine-tuning strategy that encodes spatial-temporal contextual details into prompts, enhancing LLMs' ability to reason over spatial-temporal patterns. To further improve generalization, we introduce a test time adaptation mechanism with a test data reconstructor, enabling UrbanMind to dynamically adjust to unseen test data by reconstructing LLM-generated embeddings. Extensive experiments on real-world urban datasets across multiple cities demonstrate that UrbanMind consistently outperforms state-of-the-art baselines, achieving high accuracy and robust generalization, even in zero-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。