分阶段强化学习让多机器人协同搬运月球货物更可靠
Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning

- 将搬运任务拆成举升、运输、放置三阶段,每阶段用专用策略优化
- 仿真与实测均实现全阶段稳定协作运输,硬件实验通过光学定位验证
- 适合研究多机器人协同控制或月球任务自动化部署的工程师
模块化可重构机器人系统为未来月球任务中的协同地表作业提供了可扩展解决方案。然而,由于形态依赖的拓扑变化、强负载耦合、长时决策和安全约束,协同货物运输仍具挑战。本文提出一种基于相位分解的强化学习框架,用于分布式机器人单元的协同货物运输。任务被分解为举升、运输和放置三个阶段,每个阶段通过专用联合状态策略优化,以捕捉多智能体间的耦合关系。采用集中式训练促进稳定收敛,部署时使用机载本体感知进行控制,并通过OptiTrack运动捕捉系统实现真实轨迹评估与后处理指标计算。一个基于马尔可夫状态表示的确定性相位控制器管理阶段间转换,失败敏感的同步机制确保执行过程中的协调推进与安全停机。该框架在仿真及日本宇宙航空研究开发机构(JAXA)空间探索测试设施的受控实地实验中进行了评估。结果表明,在仿真与硬件实验中所有阶段均实现了可靠的协同运输。
原文摘要 · Abstract (English)
Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative cargo transportation remains challenging due to morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. This paper proposes a phase-decomposed reinforcement learning framework for cooperative cargo transport with distributed robotic units. The task is decomposed into lifting, transportation, and placement, each optimized with a dedicated joint-state policy capturing inter-agent coupling. Centralized training promotes stable convergence, while deployment uses onboard proprioception for control and OptiTrack motion capture for ground-truth evaluation and post-processed metrics. A deterministic phase controller expressed in Markov state representation regulates transitions between stages, and a failure-sensitive synchronization mechanism ensures coordinated progression and safety-aware halting during real-world execution. The framework is evaluated in simulation and through controlled field experiments at a JAXA space exploration test facility. Results demonstrate reliable cooperative transport across all stages in both simulation and hardware experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。