arXiv:2410.09604cs.AIcs.RO2024-10被引 49

构建真实城市环境的具身智能评测平台,推动开放场景下智能体研究。

EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment

  • 基于真实城市数据构建高保真3D仿真环境,模拟行人与车辆流。
  • 设计多维度评估任务,支持智能体感知、规划与决策能力测试。
  • 提供完整接口,适用于大模型在复杂城市场景中的能力验证。

具身人工智能强调智能体身体在生成类人行为中的作用。现有研究多聚焦于封闭室内环境(如房间导航或设备操作),对开放户外场景的探索不足,主要受限于高质量仿真器、基准和数据集的缺乏。为此,本文构建了一个面向真实城市环境的具身智能评测平台。首先,基于真实城市建筑、道路等元素构建高度真实的3D仿真环境;通过结合历史数据与仿真算法,实现行人与车辆流的高保真模拟。其次,设计涵盖多种具身智能能力的评估任务。此外,提供完整的输入输出接口,使智能体可基于任务需求与环境观测进行决策并获得性能评估。该平台拓展了现有具身智能的能力边界,具备更高现实应用价值,可支撑通用人工智能的多样化应用。基于此平台,我们评估了若干主流大语言模型在不同维度与难度下的具身智能表现。

原文摘要 · Abstract (English)

Embodied artificial intelligence emphasizes the role of an agent's body in generating human-like behaviors. The recent efforts on EmbodiedAI pay a lot of attention to building up machine learning models to possess perceiving, planning, and acting abilities, thereby enabling real-time interaction with the world. However, most works focus on bounded indoor environments, such as navigation in a room or manipulating a device, with limited exploration of embodying the agents in open-world scenarios. That is, embodied intelligence in the open and outdoor environment is less explored, for which one potential reason is the lack of high-quality simulators, benchmarks, and datasets. To address it, in this paper, we construct a benchmark platform for embodied intelligence evaluation in real-world city environments. Specifically, we first construct a highly realistic 3D simulation environment based on the real buildings, roads, and other elements in a real city. In this environment, we combine historically collected data and simulation algorithms to conduct simulations of pedestrian and vehicle flows with high fidelity. Further, we designed a set of evaluation tasks covering different EmbodiedAI abilities. Moreover, we provide a complete set of input and output interfaces for access, enabling embodied agents to easily take task requirements and current environmental observations as input and then make decisions and obtain performance evaluations. On the one hand, it expands the capability of existing embodied intelligence to higher levels. On the other hand, it has a higher practical value in the real world and can support more potential applications for artificial general intelligence. Based on this platform, we evaluate some popular large language models for embodied intelligence capabilities of different dimensions and difficulties.

具身智能城市仿真评测平台大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。