通过接管数据主动学习错误,提升自动驾驶模型安全性和表现
Learning from Mistakes: Post-Training for Driving VLA with Takeover Data
- 引入接管前语言监督,让模型提前预判危险场景
- 在重构场景中进行强化微调,实现主动探索而非被动修正
- 在Bench2Drive上驾驶得分领先基线4.93,平均TTC提升11.76%
当前端到端自动驾驶的视觉-语言-动作(VLA)范式依赖静态数据集离线训练,易受分布偏移影响。现有后训练方法虽利用接管数据补充高质量专家样本,但仍存在两大缺陷:接管后监督导致安全余量有限,被动偏好优化缺乏主动探索。本文提出TakeVLA框架,通过两项互补创新克服上述问题。首先引入接管前语言监督,使VLA能主动学习错误情境下的应对策略,培养预防性思维,显著扩大安全余量。其次提出场景构想(Scenario Dreaming)强化微调机制,在重构的接管场景中鼓励主动探索。在Bench2Drive基准测试中,TakeVLA实现顶尖闭环性能,驾驶得分超越强基线SimLingo 4.93分,平均到达时间(TTC)提升11.76%。
原文摘要 · Abstract (English)
Current Vision-Language-Action (VLA) paradigms in end-to-end autonomous driving rely on offline training from static datasets, leaving them vulnerable to distribution shift. Recent post-training methods use takeover data to mitigate this by augmenting the dataset with high-quality expert takeover samples, yet they suffer from two key limitations: supervision restricted to the period after the takeover moments leads to policies with limited safety margins, and passive preference optimization lacks active exploration for optimal performance. In this paper, we propose TakeVLA, a novel VLA post-training framework that overcomes these shortcomings through two complementary innovations. First, we introduce pre-takeover language supervision, which allows the VLA to learn from mistakes proactively. By explicitly teaching the model about what to do in error-prone situations, we cultivate a precautionary mindset that anticipates hazards early and substantially enlarges safety margins. Second, we propose Scenario Dreaming, a reinforcement fine-tuning paradigm that operates in reconstruceted takeover scenarios, encouraging active exploration beyond mere preference fitting. Experiments on the Bench2Drive benchmark demonstrate that TakeVLA achieves state-of-the-art closed-loop performance, surpassing the strong VLA baseline SimLingo by 4.93 in driving score, with an enhanced safety margin as evidenced by an 11.76% increase in average TTC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。