用专家接管数据优化端到端自动驾驶,提升闭环表现。
TakeAD: Preference-based Post-optimization for End-to-end Autonomous Driving with Expert Takeover Data
- 基于专家接管数据,通过迭代模仿与偏好优化改进驾驶策略。
- 在Bench2Drive上相比纯模仿学习,显著减少系统失能次数。
- 适合研究自动驾驶安全性和闭环优化的从业者参考。
现有端到端自动驾驶方法多依赖模仿学习(IL),但存在开环训练与闭环部署间的不匹配问题,常导致驾驶员主动接管和系统失能。如何利用这些失能场景中的专家接管数据并有效扩展IL策略能力,是一个尚未充分探索的重要挑战。本文提出TakeAD,一种基于偏好的后优化框架,通过接管数据微调预训练的模仿学习策略以提升闭环驾驶性能。首先,设计了一种受真实自动驾驶系统中人类接管机制启发的高效接管数据收集流程;随后,该框架融合了迭代数据聚合(DAgger)与直接偏好优化(DPO):DAgger阶段通过模仿专家干预,使策略具备处理失能状态的基本能力;DPO阶段则进一步对齐专家在失能场景下的行为偏好。通过多轮迭代,策略逐步学习出有效的失能恢复策略,缓解开环差距。在闭环基准Bench2Drive上的实验表明,该方法优于纯模仿学习,全面消融实验验证了各组件的有效性。
原文摘要 · Abstract (English)
Existing end-to-end autonomous driving methods typically rely on imitation learning (IL) but face a key challenge: the misalignment between open-loop training and closed-loop deployment. This misalignment often triggers driver-initiated takeovers and system disengagements during closed-loop execution. How to leverage those expert takeover data from disengagement scenarios and effectively expand the IL policy's capability presents a valuable yet unexplored challenge. In this paper, we propose TakeAD, a novel preference-based post-optimization framework that fine-tunes the pre-trained IL policy with this disengagement data to enhance the closed-loop driving performance. First, we design an efficient expert takeover data collection pipeline inspired by human takeover mechanisms in real-world autonomous driving systems. Then, this post optimization framework integrates iterative Dataset Aggregation (DAgger) for imitation learning with Direct Preference Optimization (DPO) for preference alignment. The DAgger stage equips the policy with fundamental capabilities to handle disengagement states through direct imitation of expert interventions. Subsequently, the DPO stage refines the policy's behavior to better align with expert preferences in disengagement scenarios. Through multiple iterations, the policy progressively learns recovery strategies for disengagement states, thereby mitigating the open-loop gap. Experiments on the closed-loop Bench2Drive benchmark demonstrate our method's effectiveness compared with pure IL methods, with comprehensive ablations confirming the contribution of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。