提出安全探索的平衡机制,让模型与可行区域协同优化。
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
- 交替优化最大可行区域和最小不确定模型
- 零违规扩展可行区,数轮内收敛至平衡点
- 适合需要高安全性的强化学习应用
确保环境探索的安全性是强化学习中的关键问题。尽管将探索限制在可行区域内已被广泛接受为保障安全的方法,但仍有核心问题未解:通过探索可达到的最大可行区域是多少,以及如何识别该区域?本文首次揭示,安全探索的目标是找到可行区域与环境模型之间的平衡。这一结论基于两者相互依赖的特性:更大的可行区域带来更准确的环境模型,而更准确的模型又允许探索更大区域。我们提出了首个面向平衡的安全探索框架——安全均衡探索(SEE),其交替执行最大可行区域搜索与最小不确定模型构建。通过将不确定模型形式化为图结构,证明了SEE所得模型单调优化、可行区域单调扩张,且二者均收敛至安全探索的平衡点。在经典控制任务上的实验表明,该算法在零约束违反条件下成功扩展可行区域,并在数轮内实现探索平衡。
原文摘要 · Abstract (English)
Ensuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a way to ensure safety, key questions remain unresolved: what is the maximum feasible zone achievable through exploration, and how can it be identified? This paper, for the first time, answers these questions by revealing that the goal of safe exploration is to find the equilibrium between the feasible zone and the environment model. This conclusion is based on the understanding that these two components are interdependent: a larger feasible zone leads to a more accurate environment model, and a more accurate model, in turn, enables exploring a larger zone. We propose the first equilibrium-oriented safe exploration framework called safe equilibrium exploration (SEE), which alternates between finding the maximum feasible zone and the least uncertain model. Using a graph formulation of the uncertain model, we prove that the uncertain model obtained by SEE is monotonically refined, the feasible zones monotonically expand, and both converge to the equilibrium of safe exploration. Experiments on classic control tasks show that our algorithm successfully expands the feasible zones with zero constraint violation, and achieves the equilibrium of safe exploration within a few iterations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。