arXiv:2606.27861cs.RO2026-06

PPO-EAL让机器人在强化学习中更安全,精确满足物理约束且不依赖过大的惩罚系数。

PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control

论文配图:PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control
图 1 · 摘自论文原文
  • 将精确的增广拉格朗日法融入PPO,用二次惩罚项实现约束精准控制。
  • 在多个机器人任务中,安全精度和奖励表现均优于现有先进方法。
  • 适合需要高安全性与鲁棒性的实际机器人部署,如复杂装配任务。

强化学习在完成复杂机器人控制任务方面展现出巨大潜力,但多数现有方法忽视了安全要求。安全强化学习旨在最大化任务性能的同时满足明确的物理约束,然而当前算法在高效学习并精确满足约束方面仍存在困难。本文提出PPO-EAL,一种新型一阶约束策略优化框架,将精确增广拉格朗日优化集成至近端策略优化中,用于安全机器人控制。通过结合截断策略更新与精确二次惩罚项,PPO-EAL实现了理论上严谨的约束强制,无需不切实际的大惩罚因子。引入动量调节的乘子更新机制,进一步提升了对偶变量稳定性,减少了约束振荡和不安全行为,同时保持任务性能。我们在标准随机逼近假设下提供了精确性与收敛性分析。在多种基于GPU加速的机器人基准测试中——包括倒立摆平衡、双摆稳定、7自由度Franka末端执行器抓取以及四足行走——广泛验证表明,相比当前最先进的第一阶安全强化学习基线,PPO-EAL在安全精度和奖励表现上均显著领先。最后,我们展示了在接触丰富的齿轮装配任务中的零样本仿真到现实部署,结果表明PPO-EAL显著提升任务成功率,降低峰值接触力,并增强操作鲁棒性。这些成果确立了PPO-EAL作为多样安全关键型机器人系统的通用且可实用的安全强化学习框架。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as a promising solution to accomplish complex robotic control tasks; however, most of the current work ignores the safety requirements. Safe RL seeks to maximize task performance while satisfying explicit physical constraints, but current algorithms struggle to learn the policy efficiently with precise constraint satisfaction. This work proposes PPO-EAL, a novel first-order constrained policy optimization framework that integrates exact augmented Lagrangian optimization into proximal policy optimization for safe robotic control. By combining clipped policy updates with exact quadratic penalty terms, PPO-EAL achieves theoretically grounded constraint enforcement without requiring impractically large penalty factors. A momentum-regulated multiplier update further improves dual-variable stability, reducing constraint oscillation and unsafe behavior while preserving task performance. We provide exactness and convergence analysis under standard stochastic approximation assumptions. Extensive validation across diverse GPU-accelerated robotic benchmarks-including cart-pole balancing, cart-double-pendulum stabilization, 7-DoF Franka end-effector reaching, and quadrupedal locomotion-demonstrates superior safety precision and reward performance compared with state-of-the-art first-order safe RL baselines. Finally, we demonstrate zero-shot sim-to-real deployment in a contact-rich gear assembly task, where PPO-EAL substantially improves task success, reduces peak contact force, and enhances operational robustness. These results establish PPO-EAL as a general and practically deployable safe RL framework for diverse safety-critical robotic systems.

安全强化学习机器人控制约束优化策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。