arXiv:2607.20973cs.RO2026-07中稿 · the 2026 IEEE/RSJ …

用强化学习指导的控制算法,让自动驾驶赛车更有效防超车。

Deep Reinforcement-Learning-Guided Model Predictive Control for Preventing Overtakes in Autonomous Racing

论文配图:Deep Reinforcement-Learning-Guided Model Predictive Control for Preventing Overtakes in Autonomous Racing
图 1 · 摘自论文原文
  • 分层框架:强化学习生成防御参考轨迹,嵌入模型预测控制
  • 仿真中防超车时间从8.8秒提升至14.6秒,对手推进大幅减少
  • 实时性好,平均求解仅33.3毫秒,可应对高速对抗场景

本文针对自动驾驶赛车中的防守阻拦问题,提出一种基于分层强化学习引导的模型预测控制框架。与追求最快圈速不同,防守被建模为在弗雷内坐标系下对空间占用的调控。通过软演员-评论家策略层生成几何感知的防御参考轨迹,并将其作为摩擦约束下的空间正则化项嵌入非线性模型预测控制。在Thunderhill West赛道的仿真评估显示,该方法将平均防超车时间从8.8秒提升至14.6秒,同时显著抑制对手前进;车辆可利用83.4%的可用轮胎力。系统平均求解时间为33.3毫秒(标准差13.9毫秒),支持实时高速对抗交互。

原文摘要 · Abstract (English)

This paper addresses defensive blocking in autonomous racing, where a vehicle must prevent a faster opponent from overtaking while operating near its dynamic limits. Different from lap-time minimization, we formulate defense as a spatial occupancy regulation problem via a hierarchical reinforcement-learning guided model predictive control framework. A Soft Actor-Critic strategic layer operates in the Frenet domain to generate geometry-aware defensive references, which are embedded into the nonlinear model predictive control formulation as spatial regularization under friction constraints. Evaluated on the Thunderhill West circuit in simulation, the framework increases average overtake time from 8.8 s to 14.6 s while significantly reducing opponent progress. Meanwhile, it allows the vehicle to utilize 83.4% of available tire force. The framework achieves a 33.3 ms mean solve time (13.9 ms std), supporting real-time high-speed adversarial interaction.

自动驾驶强化学习模型预测控制赛车博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。