用局部反馈策略提升长序列轨迹优化效率,改善收敛性。
Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation

- 将控制序列分段并用局部反馈连接,提高采样效率。
- 在三个非线性系统上实现更快收敛与更优终端约束满足。
- 仅通过模拟即可近似系统雅可比矩阵,适合黑箱动力学。
随机单次射击轨迹优化方法(如模型预测路径积分控制,MPPI)因能处理概率性动态,在机器人领域广泛应用,尤其适用于梯度噪声大、计算成本高或不可得的情况。然而,在长动作序列下满足终端约束时,样本效率低,收敛需大量迭代。本文提出一种随机多重射击方法,通过局部反馈策略连接短控制序列,显著提升样本效率并加速收敛至目标集。此外,我们展示了仅通过模拟轨迹即可合成近似系统雅可比矩阵,使该方法适用于黑箱动力学的模型化强化学习。实验验证了该算法在三个非线性、欠驱动优化问题中的优越表现:具有解析动力学的经典倒立摆摆起任务、采用神经网络学习动力学的倒立摆摆起任务,以及执行高攻角精准失速着陆的垂直起降无人机(VTOL quadplane)任务。
原文摘要 · Abstract (English)
Stochastic single shooting trajectory optimization methods such as Model Predictive Path Integral control (MPPI) have been widely adopted in robotics due to their ability to reason about probabilistic dynamics and provide solutions where model gradients are noisy, costly to evaluate, or unavailable. However, satisfaction of terminal constraints when shooting over long action sequences is often sample inefficient, requiring a large number of iterations for convergence. In this paper, we present a stochastic multiple shooting method that optimizes short control action sequences connected via local feedback policies to improve sample efficiency and convergence to a terminal set. Additionally, we show that we are able to synthesize approximate system Jacobians purely from rollouts, making the method suitable for model-based reinforcement learning with black-box dynamics. We demonstrate the algorithm has improved sample efficiency and terminal set convergence for three nonlinear, underactuated optimization problems: a classic cartpole swingup task with analytical dynamics, a cartpole swingup task with learned neural network dynamics, and a VTOL quadplane performing a high angle-of-attack, precision post-stall landing maneuver.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。