构建真实城市导航测试集,评估智能体在东京街区的探索能力
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

- 基于602段360度视频重建东京秋叶原街区
- 175项任务中最强模型仅达人类水平的22%
- 适合研究城市级智能体导航与空间推理的学者
我们提出360CityArena,一个基于全景视频构建的逼真城市环境基准,用于评估智能体在真实城市场景中的探索能力。现有户外基准或缺乏足够写实性或复杂度,难以反映真实城市环境。360CityArena基于日本东京秋叶原街区的实景重建,涵盖85条街道、602段360度视频,包含175项人工精心设计的任务,分为环境理解、路径推理与空间推理三类,覆盖定位、地标搜索、路径规划及关系空间推理等核心能力。使用先进视觉语言模型(LMM)智能体进行评估显示,即使最强模型Gemini 2.5 Flash,在该基准上表现仅为人类水平的17.1%(人类:77.3%),揭示了城市尺度智能体导航与推理仍面临巨大挑战。360CityArena为逼真城市区域导航与空间推理提供了必要且具有挑战性的测试平台。
原文摘要 · Abstract (English)
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。