arXiv:2503.15629cs.ROcs.AI2025-03ICRA被引 2

用自监督强化学习高效训练神经控制李雅普诺夫函数,提升非线性系统稳定性验证精度。

Neural Lyapunov Function Approximation with Self-Supervised Reinforcement Learning

  • 通过自监督RL生成高质量训练数据,聚焦状态空间中误差大的区域
  • 在机器人任务中收敛速度更快,精度优于现有神经李雅普诺夫方法
  • 适合需要稳定性和安全性保障的复杂控制系统设计

传统控制李雅普诺夫函数用于设计确保系统收敛到目标状态的控制器,但对非线性系统的函数推导仍具挑战。本文提出一种新型、样本高效的神经李雅普诺夫函数近似方法,利用自监督强化学习增强训练数据生成,尤其针对状态空间中表征不准确的区域。该方法采用数据驱动的世界模型,从离策略轨迹中训练李雅普诺夫函数。在标准及目标条件化的机器人任务上进行了验证,相比当前最优的神经李雅普诺夫逼近基线,展现出更快的收敛速度和更高的近似精度。代码已开源:https://github.com/CAV-Research-Lab/SACLA.git。

原文摘要 · Abstract (English)

Control Lyapunov functions are traditionally used to design a controller which ensures convergence to a desired state, yet deriving these functions for nonlinear systems remains a complex challenge. This paper presents a novel, sample-efficient method for neural approximation of nonlinear Lyapunov functions, leveraging self-supervised Reinforcement Learning (RL) to enhance training data generation, particularly for inaccurately represented regions of the state space. The proposed approach employs a data-driven World Model to train Lyapunov functions from off-policy trajectories. The method is validated on both standard and goal-conditioned robotic tasks, demonstrating faster convergence and higher approximation accuracy compared to the state-of-the-art neural Lyapunov approximation baseline. The code is available at: https://github.com/CAV-Research-Lab/SACLA.git

控制理论强化学习稳定性分析神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。