提出安全调制强化学习方法,让无人机悬停更安全可靠
A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering
- 用安全调制器动态调整动作,确保飞行安全
- 采用分布式评论者降低价值估计偏差,提升决策可靠性
- 在真实无人机上验证,性能优于主流算法
本文提出一种无模型安全强化学习中的安全调制演员-评论者(SMAC)方法,以解决安全约束满足与Q值过估计问题。通过设计安全调制器动态调节动作,使策略可忽略安全约束而专注奖励最大化;同时提出基于分布式的评论者及其理论更新规则,有效缓解安全约束下的Q值过估计问题。在无人机悬停任务的仿真与真实场景实验中,SMAC均能有效维持安全约束,并显著优于主流基线算法。
原文摘要 · Abstract (English)
This paper proposes a safety modulator actor-critic (SMAC) method to address safety constraint and overestimation mitigation in model-free safe reinforcement learning (RL). A safety modulator is developed to satisfy safety constraints by modulating actions, allowing the policy to ignore safety constraint and focus on maximizing reward. Additionally, a distributional critic with a theoretical update rule for SMAC is proposed to mitigate the overestimation of Q-values with safety constraints. Both simulation and real-world scenarios experiments on Unmanned Aerial Vehicles (UAVs) hovering confirm that the SMAC can effectively maintain safety constraints and outperform mainstream baseline algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。