arXiv:2603.04029cs.ROcs.AI2026-03

让机器人在运行中自动适应变化,像生物一样自我改进。

Self-adapting Robotic Agents through Online Continual Reinforcement Learning with World Model Feedback

  • 用世界模型残差检测异常,触发在线持续强化学习微调。
  • 无需外部监督,通过任务表现和内部指标判断是否收敛。
  • 已在仿真四足机器人和真实模型车上验证,适合自适应系统研究者。

基于学习的机器人控制器通常离线训练、参数固定,难以应对部署过程中的意外变化。受生物启发,本文提出一种在线持续强化学习框架,实现部署期间的自动化适应。基于DreamerV3(一种基于模型的强化学习算法),该方法利用世界模型预测残差检测分布外事件,并自动触发微调。通过任务性能信号与内部训练指标共同监控适应进度,可在无外部监督和领域知识的情况下评估收敛性。该方法在多种连续控制任务中得到验证,包括高保真度仿真中的四足机器人及真实模型车辆。文中还讨论了相关指标的解读与权衡关系。结果表明,自主机器人有望突破静态训练范式,实现运行中自我反思与优化,如同生物体一般。

原文摘要 · Abstract (English)

As learning-based robotic controllers are typically trained offline and deployed with fixed parameters, their ability to cope with unforeseen changes during operation is limited. Biologically inspired, this work presents a framework for online Continual Reinforcement Learning that enables automated adaptation during deployment. Building on DreamerV3, a model-based Reinforcement Learning algorithm, the proposed method leverages world model prediction residuals to detect out-of-distribution events and automatically trigger finetuning. Adaptation progress is monitored using both task-level performance signals and internal training metrics, allowing convergence to be assessed without external supervision and domain knowledge. The approach is validated on a variety of contemporary continuous control problems, including a quadruped robot in high-fidelity simulation, and a real-world model vehicle. Relevant metrics and their interpretation are presented and discussed, as well as resulting trade-offs described. The results sketch out how autonomous robotic agents could once move beyond static training regimes toward adaptive systems capable of self-reflection and -improvement during operation, just like their biological counterparts.

强化学习机器人自适应持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。