arXiv:2503.00577cs.RO2025-03ICRA被引 1

用强化学习协作补偿模型预测控制,提升越野自动驾驶在未知地形下的性能。

Actor-Critic Cooperative Compensation to Model Predictive Control for Off-Road Autonomous Vehicles Under Unknown Dynamics

  • 采用强化学习与模型预测控制协同框架,共享预测信息实现互补。
  • 在未知松软、岩石和黏性土壤上,追踪速度误差降低最多29.2%。
  • 训练数据少于纯学习方法,且在数据不足时仍表现更优。

本研究提出一种面向未知系统动态的演员-评论家协同补偿模型预测控制器(AC3MPC),以解决高复杂度动态建模困难及实时控制可行性问题。该方法将深度强化学习与模型预测控制结合,在协同框架中使两者共享对方的预测信息,从而提升轨迹跟踪性能并保留模型预测控制的固有鲁棒性。在模拟沙质松软土、沙石混合土及黏性泥质变形土壤等未知可变形地形上进行评估,结果表明,所提控制器在纵向参考速度追踪任务中,相比纯模型基础与纯学习基础控制器分别实现最高达29.2%和10.2%的性能提升。该框架在多种未见过的地形特征下具有良好泛化能力,且所需训练数据显著少于纯学习方法,即便在训练不充分时仍保持更优表现。

原文摘要 · Abstract (English)

This study presents an Actor-Critic Cooperative Compensated Model Predictive Controller (AC3MPC) designed to address unknown system dynamics. To avoid the difficulty of modeling highly complex dynamics and ensuring realtime control feasibility and performance, this work uses deep reinforcement learning with a model predictive controller in a cooperative framework to handle unknown dynamics. The model-based controller takes on the primary role as both controllers are provided with predictive information about the other. This improves tracking performance and retention of inherent robustness of the model predictive controller. We evaluate this framework for off-road autonomous driving on unknown deformable terrains that represent sandy deformable soil, sandy and rocky soil, and cohesive clay-like deformable soil. Our findings demonstrate that our controller statistically outperforms standalone model-based and learning-based controllers by upto 29.2% and 10.2%. This framework generalized well over varied and previously unseen terrain characteristics to track longitudinal reference speeds with lower errors. Furthermore, this required significantly less training data compared to purely learning-based controller, while delivering better performance even when under-trained.

自动驾驶强化学习模型预测越野驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。