arXiv:2412.11050cs.CVcs.AI2024-12中稿 · IEEE Transactions …被引 10

用检索增强提升视觉语言模型对自动驾驶罕见场景的理解能力。

RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models

  • 引入检索增强框架,融合多模态特征与相似场景知识。
  • 在CODA-LM上达74.46分,显著优于现有方法。
  • 适合关注自动驾驶安全与可解释性的研究者。

理解并应对罕见驾驶场景对保障自动驾驶系统的安全性与可靠性至关重要。视觉语言模型(VLMs)在提升场景理解方面发挥关键作用,但面临幻觉和现实基础不足等挑战,影响其在关键驾驶场景中的表现。本文提出RAC3框架,通过频率-空间融合图像编码器、基于硬/半硬负样本挖掘的跨模态对齐训练方法,以及基于K-Means聚类与层级可导航小世界(HNSW)索引的快速查询与检索管道,增强模型性能。引入多模态思维链(CoT)提示策略,引导类比推理,减少推理过程中的幻觉。同时集成更新机制,支持持续学习。在CODA与nuScenes数据集上的大量实验表明,RAC3在多个下游任务中显著提升罕见场景理解能力。相较于先前最优方法,RAC3在CODA-LM基准上取得74.46的最高得分,并在与DriveLM等端到端框架集成时表现出一致性能提升,验证了检索增强与跨模态对齐在实现更安全、可解释自动驾驶中的有效性。

原文摘要 · Abstract (English)

Understanding and addressing corner cases is essential for ensuring the safety and reliability of autonomous driving systems. Vision-language models (VLMs) play a crucial role in enhancing scenario comprehension, yet they face significant challenges, such as hallucination and insufficient real-world grounding, which compromise their performance in critical driving scenarios. In this work, RAC3, a novel framework designed to enhance the performance of VLMs in corner case comprehension, is proposed. RAC3 integrates a frequency-spatial fusion (FSF) image encoder, a cross-modal alignment training method for embedding models with hard and semi-hard negative mining, and a fast querying and retrieval pipeline based on K-Means clustering and hierarchical navigable small world (HNSW) indexing. A multimodal chain-of-thought (CoT) prompting strategy to guide analogical reasoning and reduce hallucinations during inference is introduced. Moreover, an update mechanism is integrated into RAC3 to ensure continual learning within the framework. Extensive experiments on the CODA and nuScenes datasets demonstrate that RAC3 significantly improves corner case comprehension across multiple downstream tasks. Compared to prior state-of-the-art methods, RAC3 achieves the highest final score of 74.46 on the CODA-LM benchmark and shows consistent performance gains when integrated with end-to-end frameworks like DriveLM. These results demonstrate the effectiveness of retrieval-augmented strategies and cross-modal alignment for safer and more interpretable autonomous driving.

自动驾驶视觉语言模型检索增强罕见场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。