arXiv:2503.00287cs.RO2025-03被引 4

让机器人在复杂接触任务中更安全稳定,通过能量守恒机制提升强化学习可靠性。

Passivity-Centric Safe Reinforcement Learning for Contact-Rich Robotic Tasks

  • 引入能量约束的被动性感知训练,确保策略具备稳定性基础。
  • 部署时用被动性滤波器实时保障控制稳定,避免能量异常波动。
  • 适合高风险接触任务的机器人系统,尤其关注安全与长期运行的场景。

强化学习在各类机器人任务中取得显著进展,但在涉及高频接触的实际场景中,常忽视安全与稳定性问题。缺乏被动性保障的策略可能导致系统失稳,威胁机器人、环境及操作人员。本文揭示传统强化学习策略在接触密集任务中无法保证稳定性,并提出一种融合能量基被动控制的安全强化学习方法:在训练阶段引入基于能量的被动性约束,在部署阶段对策略输出施加被动性滤波。在接触丰富型机器人迷宫探索任务上进行对比实验,结果表明,不考虑被动性的策略虽训练时完成率高,但部署中易违反能量约束;而本方法通过被动性训练与滤波,有效保障控制稳定性并提升能量效率。真实世界实验视频及模型检查点、离线数据已公开于Hugging Face。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has achieved remarkable success in various robotic tasks; however, its deployment in real-world scenarios, particularly in contact-rich environments, often overlooks critical safety and stability aspects. Policies without passivity guarantees can result in system instability, posing risks to robots, their environments, and human operators. In this work, we investigate the limitations of traditional RL policies when deployed in contact-rich tasks and explore the combination of energy-based passive control with safe RL in both training and deployment to answer these challenges. Firstly, we reveal the discovery that standard RL policy does not satisfy stability in contact-rich scenarios. Secondly, we introduce a \textit{passivity-aware} RL policy training with energy-based constraints in our safe RL formulation. Lastly, a passivity filter is exerted on the policy output for \textit{passivity-ensured} control during deployment. We conduct comparative studies on a contact-rich robotic maze exploration task, evaluating the effects of learning passivity-aware policies and the importance of passivity-ensured control. The experiments demonstrate that a passivity-agnostic RL policy easily violates energy constraints in deployment, even though it achieves high task completion in training. The results show that our proposed approach guarantees control stability through passivity filtering and improves the energy efficiency through passivity-aware training. A video of real-world experiments is available as supplementary material. We also release the checkpoint model and offline data for pre-training at \href{https://huggingface.co/Anonymous998/passiveRL/tree/main}{Hugging Face}.

强化学习机器人控制安全性被动性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。