提出可保证测试安全的元强化学习算法,兼顾快速学习与安全约束。
Constrained Meta Reinforcement Learning with Provable Test-Time Safety
- 训练时优化策略,测试时动态修正以确保安全
- 在测试任务上实现近最优策略,样本复杂度达到理论下界
- 适合机器人、医疗等需严格安全约束的场景
元强化学习(Meta RL)使智能体能利用任务分布中的经验进行训练,从而在新测试任务上实现更快的策略学习。尽管其在降低测试任务样本复杂度方面表现优异,但现实应用如机器人和医疗领域在测试阶段往往存在安全约束。受限的元强化学习为整合安全性提供了可行框架。当前关键挑战是如何在减少样本消耗的同时,确保真实测试任务上的策略安全性。为此,我们提出一种算法,在训练中学习的策略基础上进行精炼,可在测试任务上实现近最优策略,同时具备可证明的安全性与样本复杂度保证。此外,我们推导出匹配的下界,表明该样本复杂度是紧致的。
原文摘要 · Abstract (English)
Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling faster learning of optimal policies on new test tasks. Despite its success in improving sample complexity on test tasks, many real-world applications, such as robotics and healthcare, impose safety constraints during testing. Constrained meta RL provides a promising framework for integrating safety into meta RL. An open question in constrained meta RL is how to ensure safety of the policy on the real-world test task, while reducing the sample complexity and thus, enabling faster learning of optimal policies. To address this gap, we propose an algorithm that refines policies learned during training, with provable safety and sample complexity guarantees for learning a near optimal policy on the test tasks. We further derive a matching lower bound, showing that this sample complexity is tight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。