用合成环境让多个机器人像人一样协作,安全高效完成复杂操作。
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

- 构建真实与仿真融合的合成环境,统一多机器人决策空间。
- 在多机械臂任务中实现92%以上成功率和高执行效率。
- 适合需要多智能体协同的工业自动化与服务机器人场景。
多智能体具身系统在复杂协作操作中潜力巨大,但面临空间协调、时间推理和共享工作区感知等关键挑战。受人类协作中认知规划与物理执行分离的启发,我们提出「合成环境」概念——将现实世界与仿真组件协同整合,使多个机器人能在统一决策空间中感知意图并操作。基于此,我们提出CoEnv框架,利用仿真进行安全策略探索,确保真实部署可靠性。CoEnv分三阶段运行:真实场景到仿真的重建,用于数字化物理工作区;基于视觉语言模型的动作合成,支持实时规划与代码化轨迹生成;经碰撞检测验证的仿真到真实迁移,保障安全部署。在具有挑战性的多机械臂操作基准测试中,CoEnv展现出高任务成功率和执行效率,确立了多智能体具身人工智能的新范式。
原文摘要 · Abstract (English)
Multi-agent embodied systems hold promise for complex collaborative manipulation, yet face critical challenges in spatial coordination, temporal reasoning, and shared workspace awareness. Inspired by human collaboration where cognitive planning occurs separately from physical execution, we introduce the concept of compositional environment -- a synergistic integration of real-world and simulation components that enables multiple robotic agents to perceive intentions and operate within a unified decision-making space. Building on this concept, we present CoEnv, a framework that leverages simulation for safe strategy exploration while ensuring reliable real-world deployment. CoEnv operates through three stages: real-to-sim scene reconstruction that digitizes physical workspaces, VLM-driven action synthesis supporting both real-time planning with high-level interfaces and iterative planning with code-based trajectory generation, and validated sim-to-real transfer with collision detection for safe deployment. Extensive experiments on challenging multi-arm manipulation benchmarks demonstrate CoEnv's effectiveness in achieving high task success rates and execution efficiency, establishing a new paradigm for multi-agent embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。