用视觉语言模型动态优化极端场景数据,提升自动驾驶持续学习能力。
VLM-C4L: Continual Core Dataset Learning with Corner Case Optimization via Vision-Language Models for Autonomous Driving
- 利用视觉语言模型自动筛选并增强极端场景数据
- 在Waymo和CODA数据集上实现持续学习性能稳定
- 适合需要长期适应新极端场景的自动驾驶系统
随着自动驾驶的广泛应用,应对复杂环境已成为不可避免的挑战。由于极端场景数据稀缺且多样,现有自动驾驶模型难以有效处理边缘案例,据美国国家公路交通安全管理局(NHTSA)报告,每年在美国有数百起事故由阳光眩光、雾天等边缘情况引发,导致数起致命事故。为确保系统持续稳健可靠,模型不仅需在常规场景表现良好,还需能适应新兴边缘场景,实现增量学习而不退化原有能力。然而,目前尚无方法能实现自动驾驶中可扩展的边缘案例持续学习。为此,我们提出VLM-C4L框架,通过视觉语言模型(VLMs)动态优化边缘案例数据集,并结合高质数据提取与核心数据重放策略,使模型能持续学习多样化边缘案例,同时保持对过往常规场景的性能,从而保障真实自动驾驶中的长期稳定性与适应性。我们在大规模真实驾驶数据集Waymo和边缘案例数据集CODA上进行了评估。
原文摘要 · Abstract (English)
With the widespread adoption and deployment of autonomous driving, handling complex environments has become an unavoidable challenge. Due to the scarcity and diversity of extreme scenario datasets, current autonomous driving models struggle to effectively manage corner cases. This limitation poses a significant safety risk, according to the National Highway Traffic Safety Administration (NHTSA), autonomous vehicle systems have been involved in hundreds of reported crashes annually in the United States, occurred in corner cases like sun glare and fog, which caused a few fatal accident. Furthermore, in order to consistently maintain a robust and reliable autonomous driving system, it is essential for models not only to perform well on routine scenarios but also to adapt to newly emerging scenarios, especially those corner cases that deviate from the norm. This requires a learning mechanism that incrementally integrates new knowledge without degrading previously acquired capabilities. However, to the best of our knowledge, no existing continual learning methods have been proposed to ensure consistent and scalable corner case learning in autonomous driving. To address these limitations, we propose VLM-C4L, a continual learning framework that introduces Vision-Language Models (VLMs) to dynamically optimize and enhance corner case datasets, and VLM-C4L combines VLM-guided high-quality data extraction with a core data replay strategy, enabling the model to incrementally learn from diverse corner cases while preserving performance on previously routine scenarios, thus ensuring long-term stability and adaptability in real-world autonomous driving. We evaluate VLM-C4L on large-scale real-world autonomous driving datasets, including Waymo and the corner case dataset CODA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。