用高斯图像融合提升多智能体协作的感知与控制能力
GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
- 通过分布式视觉构建全局一致的3D高斯场,动态适配各智能体视角
- 在RoboFactory上超越现有图像方法,接近点云方法性能
- 无需额外传感器,支持智能体数量扩展,适合复杂协同任务
近期,具身多智能体系统中的有效协作仍是核心挑战,尤其在需兼顾个体视角与全局环境感知的场景中。现有方法常难以平衡局部精细控制与整体场景理解,导致可扩展性差、协作质量低。本文提出GauDP,一种新颖的高斯-图像协同表示,实现多智能体协作系统的可扩展、感知感知仿学习。具体而言,GauDP从分散的RGB观测构建全局一致的3D高斯场,并动态重分配高斯属性至各智能体的本地视角,使所有智能体能自适应地查询任务关键特征,同时保持个体视角。该设计在不依赖额外传感模态(如3D点云)的前提下,实现精细控制与全局一致性。我们在包含多样化多机械臂操作任务的RoboFactory基准上评估GauDP,结果表明其性能优于现有图像基方法,接近点云驱动方法,且随智能体数量增加仍保持强可扩展性。
原文摘要 · Abstract (English)
Recently, effective coordination in embodied multi-agent systems has remained a fundamental challenge, particularly in scenarios where agents must balance individual perspectives with global environmental awareness. Existing approaches often struggle to balance fine-grained local control with comprehensive scene understanding, resulting in limited scalability and compromised collaboration quality. In this paper, we present GauDP, a novel Gaussian-image synergistic representation that facilitates scalable, perception-aware imitation learning in multi-agent collaborative systems. Specifically, GauDP constructs a globally consistent 3D Gaussian field from decentralized RGB observations, then dynamically redistributes 3D Gaussian attributes to each agent's local perspective. This enables all agents to adaptively query task-critical features from the shared scene representation while maintaining their individual viewpoints. This design facilitates both fine-grained control and globally coherent behavior without requiring additional sensing modalities (e.g., 3D point cloud). We evaluate GauDP on the RoboFactory benchmark, which includes diverse multi-arm manipulation tasks. Our method achieves superior performance over existing image-based methods and approaches the effectiveness of point-cloud-driven methods, while maintaining strong scalability as the number of agents increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。