用改进的扩散模型加速轨迹优化,减少求解时间2到20倍。
Accelerating trajectory optimization with Sobolev-trained diffusion policies

- 基于微分轨迹与反馈增益设计新型损失函数,提升扩散策略泛化性
- 仅需少量示范轨迹即可训练,使求解时间缩短2至20倍
- 适用于需要快速、高精度轨迹生成的机器人控制场景
轨迹优化(TO)通过迭代改进已知系统动力学来计算局部最优轨迹。但每次新问题都独立求解,收敛速度和结果质量依赖初始轨迹。为提高效率,可使用先前由求解器生成的轨迹训练的策略作为初始猜测进行热启动。基于扩散的策略近年成为表达力强的模仿学习模型,适合作此用途。然而,一个反直觉挑战在于:TO示范具有局部最优性,策略执行时微小偏差可能进入训练数据未覆盖的状态,导致长时程误差累积。本文聚焦于提供反馈增益的梯度型TO求解器,利用其特性,提出一种基于一阶信息的Sobolev学习损失,同时使用轨迹与反馈增益训练扩散策略。大量实验表明,该策略能避免误差累积,仅需极少示范轨迹即可学习,并将求解时间缩短2至20倍;引入一阶信息后,预测所需扩散步数更少,显著降低推理延迟。
原文摘要 · Abstract (English)
Trajectory Optimization (TO) solvers exploit known system dynamics to compute locally optimal trajectories through iterative improvements. A downside is that each new problem instance is solved independently; therefore, convergence speed and quality of the solution found depend on the initial trajectory proposed. To improve efficiency, a natural approach is to warm-start TO with initial guesses produced by a learned policy trained on trajectories previously generated by the solver. Diffusion-based policies have recently emerged as expressive imitation learning models, making them promising candidates for this role. Yet, a counterintuitive challenge comes from the local optimality of TO demonstrations: when a policy is rolled out, small non-optimal deviations may push it into situations not represented in the training data, triggering compounding errors over long horizons. In this work, we focus on learning-based warm-starting for gradient-based TO solvers that also provide feedback gains. Exploiting this specificity, we derive a first-order loss for Sobolev learning of diffusion-based policies using both trajectories and feedback gains. Through comprehensive experiments, we demonstrate that the resulting policy avoids compounding errors, and so can learn from very few trajectories to provide initial guesses reducing solving time by $2\times$ to $20 \times$. Incorporating first-order information enables predictions with fewer diffusion steps, reducing inference latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。