arXiv:2608.00113cs.RO2026-08

提出分步强化学习框架,让自动驾驶赛车学会极限漂移以最短时间过弯。

Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning

论文配图:Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning
图 1 · 摘自论文原文
  • 分阶段训练漂移控制策略,从单次过弯到完整赛道漂移
  • 仿真显示能有效降低圈速,同时稳定高侧滑状态
  • 适合研究高性能自动驾驶与极限操控的学者和工程师

在一级方程式赛车中,车手在轮胎抓地力极限内优化走线以缩短圈时;而在拉力赛中,车手故意突破牵引力进行漂移,快速调整车辆姿态以加速出弯,从而减少圈时。自主实现此类操作构成一个复杂的双重目标控制问题:既要稳定高度非线性的漂移动态,又要严格最小化圈时。为此,本文提出一种专为最短圈时漂移场景设计的规划-控制框架。首先,构建最优控制问题生成最短圈时漂移轨迹,作为先验数据训练深度强化学习漂移控制器。由于漂移涉及极大侧滑角,直接学习极具挑战,因此提出基于赛道引导的强化学习(TgRL)方法,分步推进训练:从漂移控制策略,到漂移过弯策略,最终形成完整的漂移竞速策略。奖励函数结合即时奖励与基于最短圈时目标的终局奖励。仿真结果表明,该框架使智能体成功学习到既保证车辆运动控制性能又能有效缩短圈时的漂移竞速策略。

原文摘要 · Abstract (English)

In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately reducing lap time. Autonomously executing such maneuvers formulates a complex dual-objective control problem: stabilizing highly nonlinear drift dynamics while strictly minimizing lap time. Addressing this challenge motivates the development of advanced Minimum-Lap-Time (MLT) drift control architectures. This paper proposes a planning-control framework specifically designed for MLT drifting scenario. First, we formulate an optimal control problem to generate a MLT drift planning trajectory, which is used as prior data to train a deep reinforcement learning drift controller. Given that drifting involves extremely large sideslip angles and is therefore challenging to learn directly, a Track-guided Reinforcement Learning (TgRL) drift control method is proposed to enable progressive training in a step-by-step manner, from drift control policy, to drift corner policy, and finally to a comprehensive drift race policy. The reward function incorporates both an instant reward term and an end reward term derived from the Minimum-Lap-Time objective. Simulation results demonstrate that the proposed framework enables the agent to learn a drift racing policy that not only ensures vehicle motion control performance but also effectively reduces lap time.

自动驾驶强化学习漂移控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。