arXiv:2602.02236cs.ROcs.LG2026-02

用实时循环强化学习让自动驾驶模型在线自适应,提升真实场景泛化能力。

Adaptive Control in Autonomous Driving via Real-Time Recurrent RL

  • 基于LrcSSM的实时循环强化学习,每步更新策略无需时间反向传播。
  • 在仿真与实车平台实现预训练模型的在线微调,有效应对部署时分布偏移。
  • 首次在非脉冲硬件上实现事件相机闭环控制的在线强化学习,适合机器人自适应研究者。

我们研究利用实时循环强化学习(RTRRL)对预训练自动驾驶控制策略进行在线微调,该算法内存高效,可在每个时间步直接更新策略参数而无需通过时间反向传播。我们将RTRRL扩展至支持LrcSSM——一种近期提出的非线性对角状态空间模型,并结合离线行为克隆与在线RTRRL微调,以适应部署时的分布偏移。我们在CarRacing仿真环境及配备事件相机的1:10尺度RoboRacer平台上验证了该方法,实现预训练策略在真实世界赛道跟踪中的在线微调。据我们所知,这是首次在标准(非脉冲)硬件上实现基于事件相机观测的闭环控制在线强化学习。使用LrcSSM的策略在两个场景中均表现最快且最稳定。

原文摘要 · Abstract (English)

We study online fine-tuning of pretrained control policies for autonomous driving using Real-Time Recurrent Reinforcement Learning (RTRRL), a memory-efficient algorithm that updates policy parameters at every time step without backpropagation through time. We extend RTRRL to support LrcSSM, a recently proposed nonlinear diagonal state-space model, and combine offline behavioral cloning with online RTRRL fine-tuning to adapt policies to distribution shifts at deployment. We validate the approach in the CarRacing simulation and on a 1:10-scale RoboRacer platform equipped with an event camera, where a pretrained policy is fine-tuned online during real-world line-following. To our knowledge, this is the first demonstration of online RL fine-tuning with event-camera observations on standard (non-spiking) hardware in closed-loop control. LrcSSM-based policies improve fastest and most consistently across both settings.

自动驾驶强化学习在线学习事件相机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。