arXiv:2512.13514cs.RO2025-12

用强化学习实现太空舱内机器人六自由度精准对接

Reinforcement Learning based 6-DoF Maneuvers for Microgravity Intravehicular Docking: A Simulation Study with Int-Ball2 in ISS-JEM

  • 基于PPO算法训练机器人在复杂微重力环境下的对接策略
  • 在模拟环境中实现多种干扰下稳定对接,成功率高
  • 适合研究空间机器人自主导航与真实场景迁移的团队

国际空间站内的自主飞行器在执行舱内任务时,其在传感器噪声、微小推力偏差和环境扰动下的精确对接仍是重大挑战。本文提出一种基于强化学习(RL)的六自由度(6-DoF)对接框架,针对日本实验模块(JEM)内的JAXA Int-Ball2机器人,在高保真Isaac Sim仿真环境中进行训练与评估。采用近端策略优化(PPO)算法,在动态域随机化和有界观测噪声条件下训练控制器,并显式建模推进器的气流阻力扭矩与极性结构。该设置使我们能够系统研究推进物理特性对强化学习对接性能的影响。所学策略在多种工况下均表现出稳定可靠的对接能力,为未来扩展至碰撞感知导航、安全强化学习、推进精确的仿真到现实迁移以及视觉驱动端到端对接奠定基础。

原文摘要 · Abstract (English)

Autonomous free-flyers play a critical role in intravehicular tasks aboard the International Space Station (ISS), where their precise docking under sensing noise, small actuation mismatches, and environmental variability remains a nontrivial challenge. This work presents a reinforcement learning (RL) framework for six-degree-of-freedom (6-DoF) docking of JAXA's Int-Ball2 robot inside a high-fidelity Isaac Sim model of the Japanese Experiment Module (JEM). Using Proximal Policy Optimization (PPO), we train and evaluate controllers under domain-randomized dynamics and bounded observation noise, while explicitly modeling propeller drag-torque effects and polarity structure. This enables a controlled study of how Int-Ball2's propulsion physics influence RL-based docking performance in constrained microgravity interiors. The learned policy achieves stable and reliable docking across varied conditions and lays the groundwork for future extensions pertaining to Int-Ball2 in collision-aware navigation, safe RL, propulsion-accurate sim-to-real transfer, and vision-based end-to-end docking.

强化学习空间机器人六自由度仿真对接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。