无需梯度即可从示范中学习控制策略,高效处理连续时间约束问题。
ZORMS-LfD: Learning from Demonstrations with Zeroth-Order Random Matrix Search
- 采用零阶随机矩阵搜索,不依赖梯度信息学习控制策略。
- 在连续时间无约束问题上,计算速度提升80%以上,性能相当。
- 适用于缺乏专用方法的连续时间约束问题,优于传统优化方法。
我们提出一种基于零阶随机矩阵搜索的示范学习方法(ZORMS-LfD),可在无需梯度信息的情况下,从专家示范中学习受限最优控制问题中的代价、约束与动态特性,适用于连续和离散时间系统。相比现有主流一阶方法需计算状态、控制或参数的梯度,ZORMS-LfD无需光滑损失函数假设,且突破了多数方法仅适用于离散时间的局限。在多个基准问题上,ZORMS-LfD在学习损失和计算时间上均达到或超越现有最优方法。在连续时间无约束问题中,其性能与一阶方法相当,但计算时间减少超80%;在缺乏专用方法的连续时间约束问题中,显著优于常用的无梯度优化方法Nelder-Mead。
原文摘要 · Abstract (English)
We propose Zeroth-Order Random Matrix Search for Learning from Demonstrations (ZORMS-LfD). ZORMS-LfD enables the costs, constraints, and dynamics of constrained optimal control problems, in both continuous and discrete time, to be learned from expert demonstrations without requiring smoothness of the learning-loss landscape. In contrast, existing state-of-the-art first-order methods require the existence and computation of gradients of the costs, constraints, dynamics, and learning loss with respect to states, controls and/or parameters. Most existing methods are also tailored to discrete time, with constrained problems in continuous time receiving only cursory attention. We demonstrate that ZORMS-LfD matches or surpasses the performance of state-of-the-art methods in terms of both learning loss and compute time across a variety of benchmark problems. On unconstrained continuous-time benchmark problems, ZORMS-LfD achieves similar loss performance to state-of-the-art first-order methods with an over $80$\% reduction in compute time. On constrained continuous-time benchmark problems where there is no specialized state-of-the-art method, ZORMS-LfD is shown to outperform the commonly used gradient-free Nelder-Mead optimization method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。