Lingjing让多种智能体在动态城市中协同,支持真实模拟与故障诊断。
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities

- 构建开放城市环境,统一物理引擎与任务设计
- 12个视觉语言模型在9项任务中表现受通信与资源制约
- 可追踪轨迹与通信,支持系统性失败分析
城市具身智能需要异构智能体(如无人机、地面机器人、自动驾驶车辆)在动态城市中协同。现有平台往往割裂不同具身形式,且与任务设计脱节。本文提出Lingjing,一个面向开放城市环境中异构多智能体具身智能的仿真平台。Lingjing从地理数据重建并渲染动态城市,同步多个物理引擎,并向智能体暴露共享的物理与结构化城市状态。其类似Gym的接口支持用户定义的ReAct智能体,以及单或多智能体自然语言任务,具备可配置的星型或广播通信及资源约束。每轮任务生成可归因的回放,关联智能体轨迹、通信与关系图变化、资源消耗及基于引擎的评估,实现系统性诊断。我们在共享引擎闭环协议下评估了12个视觉语言模型在9个城市任务上的表现。控制实验进一步考察了通信、可扩展性、鲁棒性与故障溯源。结果揭示了定位与长时程执行中的持续瓶颈;任务依赖的协调权衡及容量增加带来的收益递减;更重负载进一步降低成功率。Lingjing提供统一测试平台,支持可复现的端到端评估与系统性故障诊断。
原文摘要 · Abstract (English)
Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nevertheless isolate different embodiments and decouple them from task design and evaluation. We present \textbf{Lingjing}, a simulation platform for heterogeneous multi-agent embodied intelligence in open-ended urban environments. Lingjing reconstructs and renders evolving cities from geographic data, synchronizes multiple physics engines, and exposes shared physical and structured urban state to agents. Its Gym-like interface supports user-defined ReAct agents and single- or multi-agent natural-language missions with configurable star or broadcast communication and resource constraints. Each episode becomes an attribution-ready replay that links agent trajectories and communication to relation-graph changes, resource consumption, and engine-based evaluations for systematic diagnosis. We evaluate twelve vision-language models on nine urban tasks under a shared engine-in-the-loop protocol. Controlled studies further examine communication, scalability, robustness, and failure provenance. Results expose persistent bottlenecks in grounding and long-horizon execution. They also show task-dependent coordination trade-offs and diminishing returns from added capacity, while heavier workloads further reduce success. Lingjing provides a unified testbed that enables reproducible end-to-end evaluation and systematic failure diagnosis in urban multi-agent embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。