arXiv:2512.03729cs.ROcs.LG2025-12被引 2

首次在太空用强化学习控制自由飞行机器人,实现精准自主作业。

Autonomous Planning In-space Assembly Reinforcement-learning free-flYer (APIARY) International Space Station Astrobee Testing

  • 用强化学习训练六自由度控制策略,提升机器人在微重力下的自主性。
  • 在国际空间站成功验证算法,实现分钟级行为部署。
  • 适合航天机器人、深空探索与实时任务响应场景。

美国海军研究实验室(NRL)的自主空间组装强化学习自由飞行器(APIARY)实验首次在轨应用强化学习(RL)控制零重力环境中的自由飞行机器人。2025年5月27日,该团队利用国际空间站(ISS)上的NASA Astrobee机器人,首次实现了基于强化学习的自由飞行器在轨控制。采用演员-评论家结构的近端策略优化(PPO)网络,在NVIDIA Isaac Lab仿真环境中训练了鲁棒的6自由度(DOF)控制策略,通过随机化目标位姿和质量分布以增强泛化能力。本文详述了仿真测试、地面验证及在轨飞行验证全过程。此次在轨演示验证了强化学习在提升机器人自主性方面的变革潜力,可实现数分钟至数小时内的快速行为开发与部署,适用于空间探索、后勤支持及实时任务需求。

原文摘要 · Abstract (English)

The US Naval Research Laboratory's (NRL's) Autonomous Planning In-space Assembly Reinforcement-learning free-flYer (APIARY) experiment pioneers the use of reinforcement learning (RL) for control of free-flying robots in the zero-gravity (zero-G) environment of space. On Tuesday, May 27th 2025 the APIARY team conducted the first ever, to our knowledge, RL control of a free-flyer in space using the NASA Astrobee robot on-board the International Space Station (ISS). A robust 6-degrees of freedom (DOF) control policy was trained using an actor-critic Proximal Policy Optimization (PPO) network within the NVIDIA Isaac Lab simulation environment, randomizing over goal poses and mass distributions to enhance robustness. This paper details the simulation testing, ground testing, and flight validation of this experiment. This on-orbit demonstration validates the transformative potential of RL for improving robotic autonomy, enabling rapid development and deployment (in minutes to hours) of tailored behaviors for space exploration, logistics, and real-time mission needs.

强化学习空间机器人自主控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。