用学习的线性模型加速强化学习控制,实现实时运行
Accelerating Sampling-Based Control via Learned Linear Koopman Dynamics
- 用深度柯尔曼算子替代非线性动力学,提升轨迹采样效率
- 在倒立摆与四足机器人上实现接近真实模型的控制性能
- 无需解析模型,直接从数据学习,适合复杂系统实时控制
本文提出一种高效模型预测路径积分(MPPI)控制框架,适用于具有复杂非线性动力学的系统。为提升经典MPPI的计算效率并保持控制性能,我们用从交互数据中直接学习的线性深度柯尔曼算子(DKO)模型替代轨迹传播中的非线性动力学,实现更快的前向推演和更高效的轨迹采样。所提出的控制器称为MPPI-DK,在倒立摆平衡和水面车辆导航的仿真任务中进行了评估,并在四足机器人上通过参考轨迹跟踪实验完成了硬件验证。实验结果表明,MPPI-DK在控制性能上接近使用真实动力学的MPPI,同时显著降低计算开销,实现了机器人平台上的高效实时控制。
原文摘要 · Abstract (English)
This paper presents an efficient model predictive path integral (MPPI) control framework for systems with complex nonlinear dynamics. To improve the computational efficiency of classic MPPI while preserving control performance, we replace the nonlinear dynamics used for trajectory propagation with a learned linear deep Koopman operator (DKO) model, enabling faster rollout and more efficient trajectory sampling. The DKO dynamics are learned directly from interaction data, eliminating the need for analytical system models. The resulting controller, termed MPPI-DK, is evaluated in simulation on pendulum balancing and surface vehicle navigation tasks, and validated on hardware through reference-tracking experiments on a quadruped robot. Experimental results demonstrate that MPPI-DK achieves control performance close to MPPI with true dynamics while substantially reducing computational cost, enabling efficient real-time control on robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。