首个评估大模型城市时空推理能力的基准,揭示其规划与反思短板。
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
- 构建四维分解评估框架:时空理解、预测、规划、反馈反思
- 覆盖5类城市决策任务,含62,466个结构化问答对,支持细粒度诊断
- 发现先进模型在长时规划中表现不佳,需领域定制化优化
大语言模型(LLMs)在时空推理方面展现出潜力,有望用于构建支持多元城市下游应用的城市代理。然而,现有研究多聚焦结果层面指标(如预测准确率、交通效率),难以揭示其内在推理过程。为此,我们提出USTBench,首个评估LLMs作为城市代理在时空推理能力上的基准,涵盖四个分解维度:时空理解、预测、规划与反馈反思。USTBench支持五种多样化的城市决策任务和四种时空预测任务,均运行于自建交互式城市环境UAgentEnv中。基准包含62,466个结构化问答对,用于过程级评估,并提供标准化端到端任务评测,实现细粒度诊断与跨场景任务比较。通过对十三个主流LLMs的全面评估,发现尽管模型在各类城市任务中表现出潜力,但在长周期规划和动态环境中的反思适应方面仍存显著不足。值得注意的是,近期基于通用逻辑或数学问题训练的先进推理模型(如DeepSeek-R1)并未持续优于非推理型模型。这一差异凸显了提升城市时空推理能力需依赖领域专用适配方法。总体而言,USTBench为构建更自适应、高效的基于LLM的城市代理及智慧城市建设提供了坚实基础。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown emerging potential in spatiotemporal reasoning, making them promising candidates for building urban agents that support diverse urban downstream applications. Despite these benefits, existing studies primarily focus on evaluating urban LLM agent on outcome-level metrics (e.g., prediction accuracy, traffic efficiency), offering limited insight into their underlying reasoning processes. As a result, the strengths and limitations of urban LLM agents in spatiotemporal reasoning remain poorly understood. To this end, we introduce USTBench, the first benchmark to evaluate LLMs' spatiotemporal reasoning abilities as urban agents across four decomposed dimensions: spatiotemporal understanding, forecasting, planning, and reflection with feedback. Specifically, USTBench supports five diverse urban decision-making and four spatiotemporal prediction tasks, all running within our constructed interactive city environment UAgentEnv. The benchmark includes 62,466 structured QA pairs for process-level evaluation and standardized end-to-end task assessments, enabling fine-grained diagnostics and broad task-level comparison across diverse urban scenarios. Through extensive evaluation of thirteen leading LLMs, we reveal that although LLMs show promising potential across various urban downstream tasks, they still struggle in long-horizon planning and reflective adaptation in dynamic urban contexts. Notably, recent advanced reasoning models (e.g., DeepSeek-R1) trained on general logic or mathematical problems do not consistently outperform non-reasoning LLMs. This discrepancy highlights the need for domain-specialized adaptation methods to enhance urban spatiotemporal reasoning. Overall, USTBench provides a foundation to build more adaptive and effective LLM-based urban agents and broad smart city applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。