arXiv:2606.13835cs.CLcs.AI2026-06被引 2

检验大模型城市模拟中人的移动是否真实,发现看似合理实则失真。

When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation

论文配图:When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
图 1 · 摘自论文原文
  • 用真实数据对比大模型生成的移动模式,多维度验证真实性。
  • 生成结果在行程长度、停留时间等关键指标上与现实差距大。
  • 适合关注城市模拟可信度的研究者与政策制定者参考。

基于大语言模型的生成代理正被广泛用于城市模拟,但其是否能复现真实的人类移动模式仍不明确。本文提出一个验证框架,通过移动规律、时间节律、网络结构模式、语义活动转换及行为移动画像等维度,对比巴黎大区和上海的真实移动数据与模拟结果。评估发现,尽管模拟器能部分匹配高层次活动分布,但在行程长度、起止点流向、停留时长及动态转换等核心时空约束上表现不佳。此外,真实移动多样性在默认提示设置下不稳定,可能需显式的行为特征初始化。为支持可复现评估,我们还开源了支持区域级地图生成、可观测性增强仿真、移动指标计算与交通仿真的可扩展基础设施。研究强调了对大模型城市模拟进行严格实证验证的重要性,并提供了构建更真实、可复现系统的技术工具。

原文摘要 · Abstract (English)

LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives. We introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data. For this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles. Using datasets from the Greater Paris region and Shanghai, we evaluate AgentSociety and CitySim across multiple dimensions of mobility realism. Our analysis reveals a substantial gap between narrative plausibility and empirical mobility realism. Although the simulators capture some high-level semantic activity distributions, they struggle to reproduce core spatial and temporal constraints, including realistic trip-length distributions, origin-destination flows, dwell times, and transition dynamics. We further observe that realistic mobility diversity is unstable across default prompting configurations and may require explicit profile-aware initialization. To support reproducible evaluation, we also contribute scalable and open LLM-driven infrastructure for regional-scale map generation, observability-enhanced simulation, mobility-metric computation, and traffic simulation. Our findings highlight the need for rigorous empirical validation of LLM-based urban simulators and provide practical tools for building more realistic and reproducible urban simulation systems.

城市模拟大模型移动行为真实性验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。