解耦表示与协作学习,让多机器人自适应搬运不同物体更高效。
DeReCo: Decoupling Representation and Coordination Learning for Object-Adaptive Decentralized Multi-Robot Cooperative Transport
- 分三阶段训练:先集中学协作,再从局部观测重建物体表征,最后逐步去掉特权信息。
- 在三种训练物体上表现优于基线,六种未见物体均实现良好泛化。
- 适合研究多机器人协同运输、强化学习解耦训练的学者和工程师。
在去中心化执行下,跨多样化形状与物理特性的物体进行通用化多机器人协同搬运仍是根本性挑战。两个关键问题浮现:在部分可观测条件下对物体依赖的表征学习,以及在非平稳性下的多智能体强化学习(MARL)协作学习。传统方法以端到端方式联合优化物体依赖表征与协调策略,并在训练中随机化物体形状与物理属性。然而,这种联合优化导致表征与协作学习紧密耦合,引入双向干扰:部分可观测下的不准确表征会破坏协作学习,而MARL中的非平稳性进一步恶化表征学习,造成样本效率低下。为此,本文提出DeReCo,一种新型的MARL框架,通过解耦表示与协作学习,实现对象自适应的多机器人协同搬运,提升样本效率与跨对象、跨场景的泛化能力。DeReCo采用三阶段训练策略:(1) 利用特权物体信息进行集中式协作学习;(2) 从本地观测中重建物体依赖的表征;(3) 逐步去除特权信息以实现去中心化执行。该解耦机制缓解了表征与协作学习间的干扰,支持稳定且高效的训练。实验表明,DeReCo在三种训练物体的仿真环境中超越基线,在六种质量与摩擦系数各异的未见物体上实现良好泛化,并在两种未见物体的真实机器人实验中取得优异性能。
原文摘要 · Abstract (English)
Generalizing decentralized multi-robot cooperative transport across objects with diverse shapes and physical properties remains a fundamental challenge. Under decentralized execution, two key challenges arise: object-dependent representation learning under partial observability and coordination learning in multi-agent reinforcement learning (MARL) under non-stationarity. A typical approach jointly optimizes object-dependent representations and coordinated policies in an end-to-end manner while randomizing object shapes and physical properties during training. However, this joint optimization tightly couples representation and coordination learning, introducing bidirectional interference: inaccurate representations under partial observability destabilize coordination learning, while non-stationarity in MARL further degrades representation learning, resulting in sample-inefficient training. To address this structural coupling, we propose DeReCo, a novel MARL framework that decouples representation and coordination learning for object-adaptive multi-robot cooperative transport, improving sample efficiency and generalization across objects and transport scenarios. DeReCo adopts a three-stage training strategy: (1) centralized coordination learning with privileged object information, (2) reconstruction of object-dependent representations from local observations, and (3) progressive removal of privileged information for decentralized execution. This decoupling mitigates interference between representation and coordination learning and enables stable and sample-efficient training. Experimental results show that DeReCo outperforms baselines in simulation on three training objects, generalizes to six unseen objects with varying masses and friction coefficients, and achieves superior performance on two unseen objects in real-robot experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。