arXiv:2506.09800cs.RO2025-06TPAMI被引 23

让自动驾驶模型持续学习难例,兼顾通用性与安全性

Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving

  • 用强化学习动态优化难例场景,保留通用驾驶知识
  • 实测在仿真和真实数据上提升泛化能力与长程鲁棒性
  • 适合追求可扩展自动驾驶系统的研发团队

端到端自动驾驶通过学习模块整合直接将传感器输入映射为规划动作,展现出巨大潜力。然而现有基于模仿学习的模型在复杂场景下泛化能力差,部署后缺乏修正反馈机制。虽然强化学习能优化复杂场景表现,但常导致特定案例过拟合,引发灾难性遗忘和样本效率低下。为此,本文提出一种名为R2SE的新学习框架,通过持续精炼难例场景并保持通用驾驶策略,实现对模型无关端到端系统的能力提升。该框架包含三个核心组件:1)基于难例分配的通用预训练,动态识别易出错场景用于针对性优化;2)残差强化专精微调,利用强化学习优化残差修正,提升难例域性能同时保留全局驾驶知识;3)自感知适配器扩展,动态将专精策略回填至通用模型,实现持续性能改进。闭环仿真与真实数据集实验表明,R2SE在泛化能力、安全性和长时序策略鲁棒性方面均优于当前最优端到端系统,验证了强化精炼在可扩展自动驾驶中的有效性。

原文摘要 · Abstract (English)

End-to-end autonomous driving has emerged as a promising paradigm for directly mapping sensor inputs to planning maneuvers using learning-based modular integrations. However, existing imitation learning (IL)-based models suffer from generalization to hard cases, and a lack of corrective feedback loop under post-deployment. While reinforcement learning (RL) offers a potential solution to tackle hard cases with optimality, it is often hindered by overfitting to specific driving cases, resulting in catastrophic forgetting of generalizable knowledge and sample inefficiency. To overcome these challenges, we propose Reinforced Refinement with Self-aware Expansion (R2SE), a novel learning pipeline that constantly refines hard domain while keeping generalizable driving policy for model-agnostic end-to-end driving systems. Through reinforcement fine-tuning and policy expansion that facilitates continuous improvement, R2SE features three key components: 1) Generalist Pretraining with hard-case allocation trains a generalist imitation learning (IL) driving system while dynamically identifying failure-prone cases for targeted refinement; 2) Residual Reinforced Specialist Fine-tuning optimizes residual corrections using reinforcement learning (RL) to improve performance in hard case domain while preserving global driving knowledge; 3) Self-aware Adapter Expansion dynamically integrates specialist policies back into the generalist model, enhancing continuous performance improvement. Experimental results in closed-loop simulation and real-world datasets demonstrate improvements in generalization, safety, and long-horizon policy robustness over state-of-the-art E2E systems, highlighting the effectiveness of reinforce refinement for scalable autonomous driving.

自动驾驶强化学习端到端持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。