构建逼真城市仿真环境,测试机器人多模态导航与协作能力
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
- 基于虚幻引擎5生成动态城市场景,支持多机器人交互
- 提出双任务基准,评估机器人在真实城市中的感知与协作能力
- 揭示现有视觉语言模型在复杂城市环境中表现不足
近期基础模型进展推动了通用机器人在开放场景中执行多样化任务的能力,但多数研究仍局限于室内家居环境。本文提出SimWorld-Robotics(SWR),一个基于虚幻引擎5构建的大规模、逼真城市仿真平台,可程序化生成无限数量的高保真城市场景,包含行人、交通系统等动态元素,在真实度、复杂性和可扩展性上超越现有城市仿真。平台支持多机器人控制与通信。基于此,我们构建两个挑战性机器人基准任务:(1) 多模态指令跟随任务,机器人需在行人与车流干扰下依据视觉-语言指令导航至目标;(2) 多智能体搜寻任务,两名机器人需通过通信协同定位并会合。不同于已有基准,这两个任务全面评估机器人在真实场景中的关键能力,包括多模态指令理解、大空间3D推理、安全长距离导航、多机器人协作及语义化通信。实验表明,当前先进模型(如视觉-语言模型)在这些任务中表现不佳,缺乏应对城市环境所需的鲁棒感知、推理与规划能力。
原文摘要 · Abstract (English)
Recent advances in foundation models have shown promising results in developing generalist robotics that can perform diverse tasks in open-ended scenarios given multimodal inputs. However, current work has been mainly focused on indoor, household scenarios. In this work, we present SimWorld-Robotics~(SWR), a simulation platform for embodied AI in large-scale, photorealistic urban environments. Built on Unreal Engine 5, SWR procedurally generates unlimited photorealistic urban scenes populated with dynamic elements such as pedestrians and traffic systems, surpassing prior urban simulations in realism, complexity, and scalability. It also supports multi-robot control and communication. With these key features, we build two challenging robot benchmarks: (1) a multimodal instruction-following task, where a robot must follow vision-language navigation instructions to reach a destination in the presence of pedestrians and traffic; and (2) a multi-agent search task, where two robots must communicate to cooperatively locate and meet each other. Unlike existing benchmarks, these two new benchmarks comprehensively evaluate a wide range of critical robot capacities in realistic scenarios, including (1) multimodal instructions grounding, (2) 3D spatial reasoning in large environments, (3) safe, long-range navigation with people and traffic, (4) multi-robot collaboration, and (5) grounded communication. Our experimental results demonstrate that state-of-the-art models, including vision-language models (VLMs), struggle with our tasks, lacking robust perception, reasoning, and planning abilities necessary for urban environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。