用一个视觉语言动作模型实现多机器人去中心化协作,无需通信或额外对齐。
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy

- 基于预训练的VLA模型,每台机器人仅凭自身观察和身份提示独立决策。
- 在真实场景中比从零训练的去中心化模型提升64%,反应速度提高40%。
- 适合需要灵活协作的移动机器人系统,如搬运、传递物品等任务。
多机器人协作可高效完成从搬家具到建筑组装等复杂任务。然而,在移动多机器人环境中实现协调仍具挑战:集中式方法随团队规模增长而效率下降,去中心化方法通常需在推理时进行显式对齐或信息共享以应对部分可观测性。我们的核心洞察是,预训练的视觉-语言-动作(VLA)模型的具身先验可使各机器人仅依赖本地观测即可实现反应式、去中心化的协作,无需推理时的假设。我们提出CHORUS框架,将单一VLA主干适配于控制多样化多机器人团队。推理时,每台机器人独立运行一份CHORUS,仅依赖自身观测与机器人标识提示。在真实世界实验中,包括移动测距、图书馆书籍传递和洗衣篮抬升任务,CHORUS相比去中心化从零训练模型提升64个百分点,对队友行为的反应速度提升40个百分点,并超越集中式基线。结果表明,共享的VLA主干可在无需每机器人专属策略或推理时机器人间通信的情况下实现去中心化多机器人协作。
原文摘要 · Abstract (English)
Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site. However, achieving such coordination in mobile multi-robot settings remains challenging: centralized methods conditioned on the combined observations of a team scale poorly with team size, and decentralized methods that train one policy per robot often require explicit alignment procedures or information sharing at inference time to overcome partial observability. Our key insight is that the visuomotor priors of pretrained vision-language-action (VLA) models should enable reactive, decentralized collaboration from each robot's local observations alone, without these inference-time assumptions. We propose CHORUS, a framework that adapts a single VLA backbone to control diverse, multi-robot teams. At inference time, each robot runs an independent copy of CHORUS, conditioned only on its own observations and a robot-identifying prompt. In real-world experiments including mobile tape measurement, library book handovers, and laundry basket lifting, CHORUS achieves a 64% point improvement over decentralized, from-scratch models, improves reactivity to teammate behavior by 40% points, and outperforms centralized baselines. Together, these results show that a shared VLA backbone is capable of achieving decentralized multi-robot collaboration, without per-robot policies or inter-robot communication at inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。