arXiv:2607.23930eess.SYcs.RO2026-07

用可行动态映射融合强化学习与最优控制,实现安全高效机器人决策。

Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping

论文配图:Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping
图 1 · 摘自论文原文
  • 将强化学习动作映射为满足约束的实时可行参数集
  • 仿真中样本效率和闭环性能优于现有方法
  • 适合需要安全约束的实时机器人控制场景

受约束的动力系统控制需在高效完成复杂任务的同时保证递归可行性与安全性。为此,本文提出可行动作最优控制框架(FAOC),融合强化学习(RL)与最优控制(OC)。核心创新是基于优化的高效映射算法,将RL代理的静态抽象动作空间转化为状态依赖的可行参数集,严格满足动力系统的约束。该方法结合了最优控制的安全性与强化学习的灵活性。相比以往工作,无需人工设计抽象动作空间,且避免了强化学习无法保证可行性对最优控制公式的破坏。我们在机器人乒乓球实时运动规划中验证该方法,模拟实验表明,FAOC在样本效率和闭环性能上均优于当前最优基线。

原文摘要 · Abstract (English)

Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety constraints. To address these competing requirements, we present Feasible Action for Optimal Control (FAOC), a novel control framework integrating Reinforcement Learning (RL) and Optimal Control (OC). The key contribution is a computationally efficient, optimization-based mapping algorithm that transforms the RL agent's action from a static abstract set into a state-dependent feasible parameter set of the Optimal Control Problem (OCP), guaranteeing strict satisfaction of the dynamical system's constraints. Thus, FAOC effectively combines the predictable safety of OC with the flexibility of RL. In contrast to prior work, the abstract action space of the RL agent does not require expert or heuristic design, and the OCP formulation is not compromised by the inability of RL to guarantee feasibility. We apply our approach to real-time motion planning for robot table tennis, which encapsulates these challenges. Via simulated experiments, we show that FAOC outperforms state-of-the-art baselines in both sample efficiency and closed-loop performance.

强化学习最优控制机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。