arXiv:2604.19980cs.ROcs.SY2026-04

用线性升维模型提升非线性机器人控制的强化学习效率

Efficient Reinforcement Learning using Linear Koopman Dynamics for Nonlinear Robotic Systems

论文配图:Efficient Reinforcement Learning using Linear Koopman Dynamics for Nonlinear Robotic Systems
图 1 · 摘自论文原文
  • 基于科曼算子理论构建线性升维动力学模型
  • 一步预测替代多步推演,显著降低计算成本与误差
  • 在真实机械臂和四足机器人上实现高效控制

本文提出一种基于模型的强化学习框架,用于非线性机器人系统的最优闭环控制。该方法通过科曼算子理论学习线性升维动力学,并将其集成到演员-评论家架构中进行策略优化,策略表示为参数化的闭环控制器。为降低计算开销并缓解模型滚动误差,采用一步预测而非多步传播来估计策略梯度,从而实现基于流式交互数据的在线小批量策略更新。该框架在多个模拟非线性控制基准任务及两个真实硬件平台(包括Kinova Gen3机械臂和Unitree Go1四足机器人)上进行了评估。实验结果表明,相比无模型强化学习基线,样本效率显著提升;相较于基于模型的强化学习基线,控制性能更优;且控制表现可媲美依赖精确系统动力学的经典模型方法。

原文摘要 · Abstract (English)

This paper presents a model-based reinforcement learning (RL) framework for optimal closed-loop control of nonlinear robotic systems. The proposed approach learns linear lifted dynamics through Koopman operator theory and integrates the resulting model into an actor-critic architecture for policy optimization, where the policy represents a parameterized closed-loop controller. To reduce computational cost and mitigate model rollout errors, policy gradients are estimated using one-step predictions of the learned dynamics rather than multi-step propagation. This leads to an online mini-batch policy gradient framework that enables policy improvement from streamed interaction data. The proposed framework is evaluated on several simulated nonlinear control benchmarks and two real-world hardware platforms, including a Kinova Gen3 robotic arm and a Unitree Go1 quadruped. Experimental results demonstrate improved sample efficiency over model-free RL baselines, superior control performance relative to model-based RL baselines, and control performance comparable to classical model-based methods that rely on exact system dynamics.

强化学习机器人控制科曼算子模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。