用专家混合模型从观察中学习驾驶策略,提升多步预测稳定性。
Imitation Learning from Observations: An Autoregressive Mixture of Experts Approach
- 采用自回归专家混合模型拟合策略,分两阶段训练。
- 在两个真实驾驶数据集上实现高精度多步预测。
- 加入李雅普诺夫约束,确保系统长期稳定运行。
本文提出一种从观察中进行模仿学习的新方法,采用自回归专家混合模型拟合潜在策略。通过两阶段框架学习模型参数:第一阶段利用已有动力学知识估计控制输入序列,降低问题复杂度;第二阶段基于估计的控制序列,通过正则化最大似然估计学习策略。进一步引入李雅普诺夫稳定性约束,确保识别模型的渐近稳定性,从而实现准确的多步预测。在两个由人类示范采集的自动驾驶数据集上验证了该框架的有效性,证明其在建模复杂非线性动态过程中的实用性。
原文摘要 · Abstract (English)
This paper presents a novel approach to imitation learning from observations, where an autoregressive mixture of experts model is deployed to fit the underlying policy. The parameters of the model are learned via a two-stage framework. By leveraging the existing dynamics knowledge, the first stage of the framework estimates the control input sequences and hence reduces the problem complexity. At the second stage, the policy is learned by solving a regularized maximum-likelihood estimation problem using the estimated control input sequences. We further extend the learning procedure by incorporating a Lyapunov stability constraint to ensure asymptotic stability of the identified model, for accurate multi-step predictions. The effectiveness of the proposed framework is validated using two autonomous driving datasets collected from human demonstrations, demonstrating its practical applicability in modelling complex nonlinear dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。