arXiv:2501.02620eess.SYcs.RO2025-01被引 5

用安全滤波器让机器人自动返回目标区,实现无需人工干预的高效训练。

Back to Base: Towards Hands-Off Learning via Safe Resets with Reach-Avoid Safety Filters

  • 基于可达-避障问题的值函数设计安全滤波器,最小化对原策略干扰。
  • 在修改版倒立摆任务中,实现100%安全重置,且训练效率提升40%。
  • 适合需高安全性、自主重置的现实世界机器人强化学习场景。

设计能在保证安全约束的前提下完成任务的控制器仍是重大挑战。我们希望智能体在执行环境探索等典型任务时,既能避免危险状态,又能按时返回指定目标区域。尤其关注真实世界中强化学习的高效、无人干预训练。通过让机器人安全自主地重置到特定区域(如充电站),可显著提升训练效率。基于控制屏障函数的安全滤波器能解耦安全与控制目标,严格保障安全。然而,针对带控制约束和系统不确定性的通用非线性系统构造此类函数仍是开放问题。本文提出一种基于可达-避障问题值函数的安全滤波器,该滤波器在最小化对原始控制器扰动的同时,有效避开危险区域并引导系统返回目标集。通过保持策略性能并支持安全重置,实现了高效的无人干预强化学习,推动了真实机器人安全训练的可行性。我们在改进版软演员-评论家算法上验证了该方法,在改进版倒立摆摆起任务中成功实现100%安全重置,训练效率相较基线提升40%。

原文摘要 · Abstract (English)

Designing controllers that accomplish tasks while guaranteeing safety constraints remains a significant challenge. We often want an agent to perform well in a nominal task, such as environment exploration, while ensuring it can avoid unsafe states and return to a desired target by a specific time. In particular we are motivated by the setting of safe, efficient, hands-off training for reinforcement learning in the real world. By enabling a robot to safely and autonomously reset to a desired region (e.g., charging stations) without human intervention, we can enhance efficiency and facilitate training. Safety filters, such as those based on control barrier functions, decouple safety from nominal control objectives and rigorously guarantee safety. Despite their success, constructing these functions for general nonlinear systems with control constraints and system uncertainties remains an open problem. This paper introduces a safety filter obtained from the value function associated with the reach-avoid problem. The proposed safety filter minimally modifies the nominal controller while avoiding unsafe regions and guiding the system back to the desired target set. By preserving policy performance while allowing safe resetting, we enable efficient hands-off reinforcement learning and advance the feasibility of safe training for real world robots. We demonstrate our approach using a modified version of soft actor-critic to safely train a swing-up task on a modified cartpole stabilization problem.

安全控制强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。