arXiv:2503.10949cs.ROcs.AI2025-03被引 6

让机器人在真实世界中安全持续适应环境变化

Safe Continual Domain Adaptation after Sim2Real Transfer of Reinforcement Learning Policies in Robotics

  • 结合安全强化学习与持续学习,在域随机化仿真中实现部署后策略自适应
  • 实测表明策略能适应真实环境动态,且避免灾难性遗忘和安全隐患
  • 适合需要长期稳定运行的工业级机器人控制场景

领域随机化已成为强化学习中促进策略从仿真向真实机器人应用迁移的关键技术。现有方法通过大幅随机化以弥补未知系统参数,虽增强鲁棒性但导致真实世界策略效率低下。此外,预训练于域随机化仿真中的策略在部署后因强化学习优化过程固有的不稳定性及真实系统采样探索性动作可能带来的安全隐患而无法更新,限制了其对随时间变化的系统参数或环境动态的适应能力。本文提出在域随机化仿真下结合安全强化学习与持续学习,实现真实机器人控制中的部署时策略安全自适应。实验表明,该方法使策略能适配当前真实系统的领域分布与环境动态,同时最小化安全风险,并避免预训练阶段随机化仿真中出现的通用策略灾难性遗忘问题。视频与补充材料见 https://safe-cda.github.io/。

原文摘要 · Abstract (English)

Domain randomization has emerged as a fundamental technique in reinforcement learning (RL) to facilitate the transfer of policies from simulation to real-world robotic applications. Many existing domain randomization approaches have been proposed to improve robustness and sim2real transfer. These approaches rely on wide randomization ranges to compensate for the unknown actual system parameters, leading to robust but inefficient real-world policies. In addition, the policies pretrained in the domain-randomized simulation are fixed after deployment due to the inherent instability of the optimization processes based on RL and the necessity of sampling exploitative but potentially unsafe actions on the real system. This limits the adaptability of the deployed policy to the inevitably changing system parameters or environment dynamics over time. We leverage safe RL and continual learning under domain-randomized simulation to address these limitations and enable safe deployment-time policy adaptation in real-world robot control. The experiments show that our method enables the policy to adapt and fit to the current domain distribution and environment dynamics of the real system while minimizing safety risks and avoiding issues like catastrophic forgetting of the general policy found in randomized simulation during the pretraining phase. Videos and supplementary material are available at https://safe-cda.github.io/.

机器人控制安全强化学习持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。