arXiv:2607.24320cs.ROcs.SY2026-07

用持续强化学习让赛车在新赛道快速适应,15分钟内超越传统控制方法。

Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

论文配图:Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform
图 1 · 摘自论文原文
  • 基于持续反向传播的强化学习框架,仅用真实数据训练通用策略。
  • 在15分钟内微调后性能超过经典控制器,实现快速适应。
  • 适合需要实时自适应的机器人控制场景,如自动驾驶、竞赛机器人。

现代机器人面临环境变化的挑战,尤其当仿真无法覆盖所有真实场景时,物理世界中的强化学习变得必要。持续强化学习为此提供解决方案,但相关框架与方法仍不充分。自主赛车,特别是RoboRacer竞赛,为该问题提供了理想测试平台:在新赛道-地面组合上以最少的新经验学会驾驶,天然构成持续学习任务。本文提出一种基于持续反向传播的持续强化学习框架,仅使用真实世界数据,先训练出通用策略,再在15分钟内微调,即可超越经典控制器。此外,还提出了基于离线强化学习的对比方法,并对两种方法的可塑性进行了仿真分析。

原文摘要 · Abstract (English)

A key challenge in modern robotics is to adapt to changing environments, a challenge that is exacerbated when simulations cannot encompass every possible real-world configuration, and therefore Reinforcement Learning (RL) in the physical world becomes necessary. Continual Reinforcement Learning provides the tools to address this challenge; however, both the frameworks and the methods remain underexplored. Autonomous Racing and in particular the RoboRacer competition provide a testing ground for such methods, as learning to drive on a new track-floor combination with the least amount of new experience naturally frames a continual learning problem. This work tries to address this gap by proposing a continual RL framework based on Continual Backpropagation that is able, with only real-world data, to train a generalistic policy on a set of tracks and then fine- tune it within 15 minutes to outperform classical controllers. Furthermore, a comparison method based on offline RL is proposed, and a simulation analysis of the plasticity properties of the methods is conducted.

持续学习强化学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。