arXiv:2506.02767cs.LGcs.RO2025-06

用快速轨迹优化加速模型强化学习,提升训练效率与成功率。

Accelerating Model-Based Reinforcement Learning using Non-Linear Trajectory Optimization

  • 结合iLQR生成探索性轨迹初始化策略,加速收敛
  • 在小车摆杆任务中减少45.9%执行时间,成功率100%
  • 适合追求高效训练的模型强化学习研究者

本文针对最先进的模型强化学习算法MC-PILCO在策略优化中收敛慢的问题,将其与适用于非线性系统的快速轨迹优化方法iLQR结合,提出探索增强型MC-PILCO(EB-MC-PILCO)。该方法利用iLQR生成信息丰富且具探索性的轨迹来初始化策略,显著减少优化步骤。在小车摆杆任务上的实验表明,相较于标准MC-PILCO,EB-MC-PILCO在四次试验内完成任务时,执行时间最多缩短45.9%;即使在MC-PILCO迭代次数更少的情况下,仍能保持100%的成功率并更快求解。

原文摘要 · Abstract (English)

This paper addresses the slow policy optimization convergence of Monte Carlo Probabilistic Inference for Learning Control (MC-PILCO), a state-of-the-art model-based reinforcement learning (MBRL) algorithm, by integrating it with iterative Linear Quadratic Regulator (iLQR), a fast trajectory optimization method suitable for nonlinear systems. The proposed method, Exploration-Boosted MC-PILCO (EB-MC-PILCO), leverages iLQR to generate informative, exploratory trajectories and initialize the policy, significantly reducing the number of required optimization steps. Experiments on the cart-pole task demonstrate that EB-MC-PILCO accelerates convergence compared to standard MC-PILCO, achieving up to $\bm{45.9\%}$ reduction in execution time when both methods solve the task in four trials. EB-MC-PILCO also maintains a $\bm{100\%}$ success rate across trials while solving the task faster, even in cases where MC-PILCO converges in fewer iterations.

强化学习轨迹优化模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。