arXiv:2409.12045cs.LGcs.RO2024-09CoRL被引 5

提出可学习约束的ATACOM方法,提升强化学习长期安全性与不确定性处理能力。

Handling Long-Term Safety and Uncertainty in Safe Reinforcement Learning

  • 在ATACOM基础上引入可学习约束,融合先验知识与数据驱动策略
  • 训练中保持更安全行为,最终性能优于或媲美现有先进方法
  • 适合需长期安全保证的机器人实机部署场景

安全是制约强化学习在真实机器人中应用的关键问题。尽管多数安全强化学习方法无需约束和运动学先验知识,仅依赖数据,但在复杂现实场景中部署困难。相反,基于模型的方法通过将约束与动力学先验融入学习框架,已证明可在真实机器人上直接运行。然而,尽管机器人动力学的近似模型通常可得,安全约束却因任务而异且难以获取:可能过于复杂无法解析表达,计算成本过高,或难以预先设想长期安全需求。本文通过扩展安全探索方法ATACOM,引入可学习约束,重点解决长期安全与不确定性处理问题。所提方法在最终性能上达到或超越当前最优水平,同时训练过程保持更安全的行为表现。

原文摘要 · Abstract (English)

Safety is one of the key issues preventing the deployment of reinforcement learning techniques in real-world robots. While most approaches in the Safe Reinforcement Learning area do not require prior knowledge of constraints and robot kinematics and rely solely on data, it is often difficult to deploy them in complex real-world settings. Instead, model-based approaches that incorporate prior knowledge of the constraints and dynamics into the learning framework have proven capable of deploying the learning algorithm directly on the real robot. Unfortunately, while an approximated model of the robot dynamics is often available, the safety constraints are task-specific and hard to obtain: they may be too complicated to encode analytically, too expensive to compute, or it may be difficult to envision a priori the long-term safety requirements. In this paper, we bridge this gap by extending the safe exploration method, ATACOM, with learnable constraints, with a particular focus on ensuring long-term safety and handling of uncertainty. Our approach is competitive or superior to state-of-the-art methods in final performance while maintaining safer behavior during training.

安全强化学习长期安全可学习约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。