arXiv:2608.06481cs.ROcs.AI2026-08

用李雅普诺夫分析指导进化优化,让仿真策略更安全可靠地迁移到现实世界。

LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning

论文配图:LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning
图 1 · 摘自论文原文
  • 结合李雅普诺夫稳定性分析与统计模型检验,迭代优化并验证策略安全性。
  • 在双杆倒立和3D四旋翼上实现安全可靠的仿真到现实迁移。
  • 适合需要高安全性的机器人控制场景,如自动驾驶、工业自动化。

在仿真中训练安全且鲁棒的控制器,并系统性评估其在真实世界部署的准备度,仍是仿真到现实迁移中的关键挑战。为此,我们提出 LyEvO,一种融合约束进化优化与基于统计模型检验(SMC)的验证方法,并引入李雅普诺夫(Lyapunov)稳定性分析的物理引导框架。利用系统动力学先验知识,LyEvO 通过李雅普诺夫分析计算初始候选稳定区域。随后,在该区域内迭代采样运行场景,联合优化并统计验证策略,并根据验证结果扩展稳定区域边界。这一集成流程提供了实际可用的部署就绪判定标准。我们在双杆倒立(Cartpole)和3D四旋翼(3D Quadrotor)基准上通过大量仿真与针对性实机实验评估了 LyEvO,验证了其在仿真到现实迁移中的安全性与鲁棒性。

原文摘要 · Abstract (English)

Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain key challenges in sim-to-real transfer. To address this, we propose LyEvO, a physics-grounded framework that combines constrained Evolutionary Optimization and Statistical Model Checking (SMC)-based verification with Lyapunov-based stability analysis. Leveraging prior knowledge of the system dynamics, LyEvO uses Lyapunov analysis to compute an initial candidate stability region. An iterative loop then uses operational scenarios drawn from this region to jointly optimize and statistically verify a policy, and subsequently expands the region's boundaries based on the verification outcome. This integrated procedure provides a practical criterion for assessing deployment readiness. We evaluate LyEvO on Cartpole and 3D Quadrotor benchmarks through extensive simulations and targeted real-world experiments, demonstrating safe and robust sim-to-real transfer.

安全控制仿真迁移进化优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。