构建城市级空间智能测试平台,推动AI理解真实城市环境。
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

- 采集18条实测轨迹,每条平均83.7公里,覆盖真实城市复杂场景。
- 建立城市适配重建基线,实现可闭环运行的数字孪生仿真系统。
- 聚焦可扩展性、外推能力与不确定性,助力AI空间认知研究。
人类能导航陌生城市并逐步形成覆盖数十平方公里的空间认知地图。人工智能能否在类似尺度上构建空间表征?尽管近期基础模型在场景重建与具身智能方面取得进展,但扩展至整个城市仍是开放挑战,主要受限于缺乏城市尺度数据。为此,我们推出WildCity,一个由自动驾驶车队在复杂城市环境中采集的真实世界多模态数据集。数据集包含18条轨迹,每条平均长度83.7公里,保留了野外感知的核心挑战,如动态物体、光照变化和不完美相机位姿。我们进一步建立了城市定制的重建基线,并将重建环境转化为闭环仿真器。除数据集与基线外,我们系统分析了实现仿真级城市数字孪生的关键挑战:可扩展性、外推能力和不确定性。最终,WildCity旨在推动城市级渲染的进步,并更广泛地促进具有类人空间感知、记忆与推理能力的AI发展。
原文摘要 · Abstract (English)
Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models have advanced scene reconstruction and embodied intelligence, scaling to entire cities remains an open challenge, primarily due to the lack of city-scale data. To bridge the gap, we introduce WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments. Our dataset includes 18 trajectories, each averaging 83.7 kilometers in length, and preserves the core challenges of in-the-wild perception, e.g., dynamic objects, lighting variations, and imperfect camera poses. We further establish an urban-tailored reconstruction baseline and convert the reconstructed environments into a closed-loop simulator. Beyond the dataset and baseline, we systematically analyze the key challenges on the path to simulation-ready urban digital twins: scalability, extrapolation, and uncertainty. Ultimately, WildCity aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition. Project page: https://han-xiangyu.github.io/Wild-City/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。