用逆强化学习学场景动态,让自动驾驶车更准预测轨迹。
Inverse RL Scene Dynamics Learning for Nonlinear Predictive Control in Autonomous Vehicles
- 用深度网络学场景动态,替代传统车辆模型
- 在虚拟、实车和公开道路均优于基线方法
- 适合做自动驾驶决策与控制的科研与工程人员
本文提出基于深度学习的非线性模型预测控制器(DL-NMPC-SD),用于自主导航。该方法结合先验车辆模型与从时序测距信息中学习到的场景动态模型,后者负责估计期望轨迹并修正实际系统模型。通过将场景动态模型嵌入深层神经网络,实现对高阶运行状态空间的非线性逼近。模型利用包含增强记忆组件的时序测距观测与系统状态联合训练。采用逆强化学习与贝尔曼最优性原理,基于改进的深度Q学习算法进行端到端训练,以估计最优动作价值函数形式的期望状态轨迹。在三个实验中评估:一是在GridSim虚拟环境;二是在RovisLab AMTU平台完成室内外导航任务;三是全尺寸自动驾驶测试车在公共道路上行驶。结果表明,该方法在各项指标上优于动态窗口法(DWA)以及两种先进的端到端与强化学习方法。
原文摘要 · Abstract (English)
This paper introduces the Deep Learning-based Nonlinear Model Predictive Controller with Scene Dynamics (DL-NMPC-SD) method for autonomous navigation. DL-NMPC-SD uses an a-priori nominal vehicle model in combination with a scene dynamics model learned from temporal range sensing information. The scene dynamics model is responsible for estimating the desired vehicle trajectory, as well as to adjust the true system model used by the underlying model predictive controller. We propose to encode the scene dynamics model within the layers of a deep neural network, which acts as a nonlinear approximator for the high order state-space of the operating conditions. The model is learned based on temporal sequences of range sensing observations and system states, both integrated by an Augmented Memory component. We use Inverse Reinforcement Learning and the Bellman optimality principle to train our learning controller with a modified version of the Deep Q-Learning algorithm, enabling us to estimate the desired state trajectory as an optimal action-value function. We have evaluated DL-NMPC-SD against the baseline Dynamic Window Approach (DWA), as well as against two state-of-the-art End2End and reinforcement learning methods, respectively. The performance has been measured in three experiments: i) in our GridSim virtual environment, ii) on indoor and outdoor navigation tasks using our RovisLab AMTU (Autonomous Mobile Test Unit) platform and iii) on a full scale autonomous test vehicle driving on public roads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。