arXiv:2608.23452cs.ROcs.AI2026-08中稿 · publication at the…

无奖励信号下,让太空机器人持续自适应故障环境。

Reward-Free Continual Adaptation for Resilient Space Robots

论文配图:Reward-Free Continual Adaptation for Resilient Space Robots
图 1 · 摘自论文原文
  • 用隐状态世界模型预训练,预测奖励结构。
  • 部署后仅更新动态模型,无需新奖励信号。
  • 适合硬件退化严重的深空任务机器人。

太空机器人在极端环境中运行,硬件退化会严重威胁传统控制策略。尽管持续强化学习可实现在线自适应,但其通常依赖部署时的奖励信号,而太空环境下因缺乏外部跟踪系统,精确奖励计算往往不可行。为应对不可观测奖励的问题,我们提出一种无奖励的持续学习框架,利用隐状态世界模型。通过在多样化仿真中预训练基于模型的智能体,世界模型学会在隐空间中预测奖励结构。部署至严重硬件退化环境后,冻结观察编码器与奖励预测器,仅通过无监督回放更新世界模型的转移动态。随后,智能体完全基于该更新后的世界模型生成的想象轨迹进行策略训练,实现对动态变化的适应,而无需接收新奖励。我们在模拟的行星穿越、轨道导航和精密装配任务中,针对严重形态故障验证了该方法的有效性。

原文摘要 · Abstract (English)

Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access to a reward signal during deployment. However, precise reward computation in space is often infeasible due to the lack of external tracking systems and the overall complexity of the environment. To address the challenge of unobservable rewards, we introduce a reward-free continual learning framework that leverages latent-state world models. By pre-training a model-based agent across diverse simulations, the world model learns a robust predictor of the reward structure within its latent space. Upon deployment to an environment with severe hardware degradation, we freeze the observation encoder and reward predictor to update only the transition dynamics of the world model through unsupervised rollouts. By training the policy entirely on imagined trajectories generated by this updated world model, the agent adapts to altered dynamics without receiving new rewards. We demonstrate our approach across simulated planetary traversal, orbital navigation, and precision assembly tasks subjected to severe morphological failures.

持续学习强化学习太空机器人无奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。