arXiv:2510.20347cs.ROcs.MA2025-10中稿 · IEEE iSpaRo 2025被引 2

模块化月球机器人用分散强化学习实现零样本自适应配置

Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots

  • 每个模块独立学策略,轮子用SAC,机械臂用PPO
  • 转向误差仅3.63°,抓取成功率84.6%,电机扭矩降95.4%
  • 无需重新训练即可适配新结构,适合月面任务机器人

模块化可重构机器人适用于特定任务的太空作业,但形态组合的指数增长阻碍了统一控制。本文提出一种去中心化强化学习(Dec-RL)方案:每个模块独立学习自身策略——轮式模块采用Soft Actor-Critic(SAC)进行运动控制,7自由度机械臂使用Proximal Policy Optimization(PPO)完成转向与操作,实现了对未见过配置的零样本泛化。仿真中,转向策略的期望与实际角度均值绝对误差为3.63°;操作策略在目标偏移标准下成功率达84.6%;轮子策略相较基线平均电机扭矩降低95.4%,同时保持99.6%的成功率。月面类比实地测试验证了零样本集成在自主移动、转向及初步重构对齐上的可行性。系统在同步、并行、串行三种策略执行模式间平滑切换,无空闲状态或控制冲突,展现出可扩展、可复用且鲁棒的模块化月球机器人控制范式。

原文摘要 · Abstract (English)

Modular reconfigurable robots suit task-specific space operations, but the combinatorial growth of morphologies hinders unified control. We propose a decentralized reinforcement learning (Dec-RL) scheme where each module learns its own policy: wheel modules use Soft Actor-Critic (SAC) for locomotion and 7-DoF limbs use Proximal Policy Optimization (PPO) for steering and manipulation, enabling zero-shot generalization to unseen configurations. In simulation, the steering policy achieved a mean absolute error of 3.63° between desired and induced angles; the manipulation policy plateaued at 84.6 % success on a target-offset criterion; and the wheel policy cut average motor torque by 95.4 % relative to baseline while maintaining 99.6 % success. Lunar-analogue field tests validated zero-shot integration for autonomous locomotion, steering, and preliminary alignment for reconfiguration. The system transitioned smoothly among synchronous, parallel, and sequential modes for Policy Execution, without idle states or control conflicts, indicating a scalable, reusable, and robust approach for modular lunar robots.

机器人控制强化学习模块化月面任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。