arXiv:2602.15817cs.LGcs.RO2026-02被引 1

用强化学习解决未知可行性的安全控制问题,自动找能保安全的初始状态范围。

Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning

  • 设计新方法同时探索可行初始状态并学安全策略
  • 在两个模拟器上覆盖范围比现有最佳方法多50%以上
  • 适合需要高安全鲁棒性的复杂控制系统研究

深度强化学习在高维控制任务中表现优异,但用于可达性问题时存在根本矛盾:可达性目标是最大化系统长期保持安全的状态集合,而强化学习则优化特定分布下的期望回报。这可能导致对低概率但仍在安全集内的状态表现不佳。传统鲁棒优化虽可考虑多种初始条件,但其解的存在依赖于未知的可行性。本文提出可行性引导探索(FGE)方法,同时识别出存在安全策略的可行初始条件子集,并在此基础上学习安全策略。实验表明,在MuJoCo和Kinetix模拟器中,使用像素观测的挑战性任务下,FGE所学策略的覆盖范围超过现有最佳方法50%以上。

原文摘要 · Abstract (English)

Recent advances in deep reinforcement learning (RL) have achieved strong results on high-dimensional control tasks, but applying RL to reachability problems raises a fundamental mismatch: reachability seeks to maximize the set of states from which a system remains safe indefinitely, while RL optimizes expected returns over a user-specified distribution. This mismatch can result in policies that perform poorly on low-probability states that are still within the safe set. A natural alternative is to frame the problem as a robust optimization over a set of initial conditions that specify the initial state, dynamics and safe set, but whether this problem has a solution depends on the feasibility of the specified set, which is unknown a priori. We propose Feasibility-Guided Exploration (FGE), a method that simultaneously identifies a subset of feasible initial conditions under which a safe policy exists, and learns a policy to solve the reachability problem over this set of initial conditions. Empirical results demonstrate that FGE learns policies with over 50% more coverage than the best existing method for challenging initial conditions across tasks in the MuJoCo simulator and the Kinetix simulator with pixel observations.

强化学习安全控制鲁棒优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。