arXiv:2505.07321cs.RO2025-05被引 2

无需仿真预训练,车载强化学习让赛车实时加速,20分钟内提升竞速性能。

Drive Fast, Learn Faster: On-Board RL for High Performance Autonomous Racing

  • 采用残差强化学习结构,结合多步时序差分与异步训练,提升决策效率。
  • 在F1TENTH平台实测,20分钟训练即实现11.5%的圈速优化,超越当前最佳方法。
  • 支持端到端训练,适合高动态实时自主系统如自动驾驶赛车场景。

自动驾驶赛车因非线性动力学、高速运动及动态环境下的实时决策需求而面临独特挑战。传统强化学习方法依赖大量仿真预训练,难以有效迁移到真实世界。本文提出一种鲁棒的车载强化学习框架,无需仿真预训练即可实现真实环境直接适配。系统引入改进的Soft Actor-Critic(SAC)算法,通过残差强化学习结构,在实时运行中增强经典控制器,融合多步时序差分(TD)学习、异步训练流水线与启发式延迟奖励调整(HDRA),显著提升样本效率与训练稳定性。在F1TENTH赛车平台上广泛验证,残差强化学习控制器持续优于基线,相比现有最先进方法实现最高11.5%的圈速降低,仅需20分钟训练。此外,无基线控制器的端到端强化学习控制器亦超越此前最佳结果,支持持续赛道学习。该框架为高性能自动驾驶赛车提供了可靠解决方案,并为其他实时动态自主系统提供新方向。

原文摘要 · Abstract (English)

Autonomous racing presents unique challenges due to its non-linear dynamics, the high speed involved, and the critical need for real-time decision-making under dynamic and unpredictable conditions. Most traditional Reinforcement Learning (RL) approaches rely on extensive simulation-based pre-training, which faces crucial challenges in transfer effectively to real-world environments. This paper introduces a robust on-board RL framework for autonomous racing, designed to eliminate the dependency on simulation-based pre-training enabling direct real-world adaptation. The proposed system introduces a refined Soft Actor-Critic (SAC) algorithm, leveraging a residual RL structure to enhance classical controllers in real-time by integrating multi-step Temporal-Difference (TD) learning, an asynchronous training pipeline, and Heuristic Delayed Reward Adjustment (HDRA) to improve sample efficiency and training stability. The framework is validated through extensive experiments on the F1TENTH racing platform, where the residual RL controller consistently outperforms the baseline controllers and achieves up to an 11.5 % reduction in lap times compared to the State-of-the-Art (SotA) with only 20 min of training. Additionally, an End-to-End (E2E) RL controller trained without a baseline controller surpasses the previous best results with sustained on-track learning. These findings position the framework as a robust solution for high-performance autonomous racing and a promising direction for other real-time, dynamic autonomous systems.

强化学习自动驾驶赛车实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。