arXiv:2608.24525cs.RO2026-08

用滚动预测生成高质量专家数据,提升端到端自动驾驶安全性。

RoG-DAgger: Rollout-Guided Post-Training for End-to-End Driving

论文配图:RoG-DAgger: Rollout-Guided Post-Training for End-to-End Driving
图 1 · 摘自论文原文
  • 通过短时运动预测构建安全关键状态下的专家示范
  • 在Bench2Drive上使驾驶得分提升5.3点,成功率提高6.2个百分点
  • 适合追求高鲁棒性的自动驾驶系统后训练优化

近期端到端自动驾驶系统在闭环基准测试中表现优异,但仍主要依赖固定专家数据进行开环模仿学习。这种训练-推理不一致导致策略在策略诱导状态中易出错,累积误差可能引发安全事故。现有驾驶DAgger流程面临三大挑战:专家解空间受限、接管时机不当、专家决策依赖学生不可见信息。为此,我们提出RoG-DAgger,利用短时运动学滚动预测,在安全临界状态下构建高质量专家示范。该方法扩展了轨迹与速度解空间,通过滚动评估候选方案以提供预防性监督;利用滚动可解性判定接近“不可逆点”的接管时机;对齐专家视野与学生感知,实现兼容性监督。在分布内(含长时程)和分布外评估中,RoG-DAgger使SimLingo模型在Bench2Drive上驾驶得分提升5.3点,成功率提高6.2个百分点;在Longest6 v2上驾驶得分从22翻倍至44;在Fail2Drive上分布外成功率由55%提升至66%。

原文摘要 · Abstract (English)

Recent end-to-end driving systems demonstrate strong performance on closed-loop benchmarks, yet are still predominantly trained on fixed expert-collected data using open-loop imitation learning. This training-inference mismatch leaves the policy vulnerable in policy-induced states, where accumulated errors can lead to safety-critical failures. A promising post-training approach to overcome this issue is Dataset Aggregation (DAgger), which gathers expert demonstrations in policy-induced states and subsequently fine-tunes the policy on the resulting aggregated dataset. Existing driving DAgger pipelines, however, face three challenges: i) the expert is restricted to a limited trajectory-and-speed solution space, ii) takeover may occur too early or too late relative to impending failures, and iii) privileged expert decisions may rely on information unavailable to the student. To address this, we introduce RoG-DAgger, a post-training framework that uses short-horizon kinematic rollouts to construct high-quality expert demonstrations in safety-critical states. Specifically, RoG-DAgger expands the expert's trajectory-and-speed solution space and evaluates candidate plans through rollout to construct preventive supervision. Moreover, it uses rollout solvability to time the takeover near the estimated point of no return. Lastly, it aligns the expert's field of view with that of the student to provide student-compatible supervision. Across in-distribution (including long-horizon) and out-of-distribution evaluations, RoG-DAgger improves the end-to-end model SimLingo by 5.3 driving-score points and 6.2 percentage points in success rate on Bench2Drive, doubles its driving score from 22 to 44 on Longest6 v2, and improves out-of-distribution success rate from 55\% to 66\% on Fail2Drive.

自动驾驶后训练强化学习安全增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。