arXiv:2503.05092cs.ROcs.AI2025-03ICRA被引 3

用抽象模拟器训练多机器人协作策略,成功部署到真实机器人上。

Multi-Robot Collaboration through Reinforcement Learning and Abstract Simulation

  • 用高阶抽象模拟器替代真实物理模拟,降低训练成本。
  • 通过三类改进使抽象模型策略成功迁移到真实机器人,表现接近竞赛级非学习方法。
  • 适合研究低成本、可迁移的多机器人协同控制方案的开发者。

人类团队通过构建世界与代理动态的抽象心理模型来协作完成复杂任务。与当前多数机器人学习依赖高保真模拟器和强化学习(RL)不同,本文探究抽象模拟器在多智能体强化学习(MARL)中的应用潜力及其策略在物理机器人上的部署效果。抽象模拟器以高层抽象建模目标任务,忽略影响决策的大量环境细节。策略在抽象环境中训练后,通过独立获取的低层感知与运动控制模块部署到物理机器人。我们识别出三类关键改进:模拟器保真度提升、训练优化与模拟随机性。在合作机器人足球任务中进行大规模消融实验,验证各类改进对策略迁移的有效性。结果表明,本方法生成的策略性能与年度RoboCup竞赛中经调优的非学习行为架构相当。总体证明,使用高度抽象的世界模型也可通过MARL训练出可用于真实机器人的协作行为。

原文摘要 · Abstract (English)

Teams of people coordinate to perform complex tasks by forming abstract mental models of world and agent dynamics. The use of abstract models contrasts with much recent work in robot learning that uses a high-fidelity simulator and reinforcement learning (RL) to obtain policies for physical robots. Motivated by this difference, we investigate the extent to which so-called abstract simulators can be used for multi-agent reinforcement learning (MARL) and the resulting policies successfully deployed on teams of physical robots. An abstract simulator models the robot's target task at a high-level of abstraction and discards many details of the world that could impact optimal decision-making. Policies are trained in an abstract simulator then transferred to the physical robot by making use of separately-obtained low-level perception and motion control modules. We identify three key categories of modifications to the abstract simulator that enable policy transfer to physical robots: simulation fidelity enhancements, training optimizations and simulation stochasticity. We then run an empirical study with extensive ablations to determine the value of each modification category for enabling policy transfer in cooperative robot soccer tasks. We also compare the performance of policies produced by our method with a well-tuned non-learning-based behavior architecture from the annual RoboCup competition and find that our approach leads to a similar level of performance. Broadly we show that MARL can be use to train cooperative physical robot behaviors using highly abstract models of the world.

多机器人强化学习抽象模拟策略迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。