arXiv:2503.13861cs.CVcs.AI2025-03CVPR被引 24

用检索增强提升视觉语言模型的自动驾驶决策能力

RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving

  • 通过三阶段检索生成流程增强视觉语言模型决策
  • 在自建努斯尼斯数据集上准确率与F1显著提升
  • 适合研究自动驾驶高层决策与多模态模型融合者

准确理解并决策高层元动作对保障自动驾驶系统可靠性和安全性至关重要。尽管视觉语言模型(VLMs)在多种自动驾驶任务中展现出巨大潜力,但常受限于空间感知不足和幻觉问题,导致在复杂场景下表现不佳。为此,本文提出一种检索增强决策框架(RAD),旨在提升VLM在自动驾驶场景中生成元动作的可靠性。RAD采用检索增强生成(RAG)管道,通过嵌入流、检索流和生成流三个阶段动态优化决策准确性。此外,我们基于NuScenes数据集构建特定数据集,对VLM进行微调,以增强其空间感知与鸟瞰图理解能力。在自建的基于NuScenes的数据集上的大量实验表明,RAD在匹配准确率、F1分数及自定义综合评分等多项关键指标上均优于基线方法,验证了其在自动驾驶元动作决策中的有效性。

原文摘要 · Abstract (English)

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autonomous driving tasks, they often suffer from limitations such as inadequate spatial perception and hallucination, reducing their effectiveness in complex autonomous driving scenarios. To address these challenges, we propose a retrieval-augmented decision-making (RAD) framework, a novel architecture designed to enhance VLMs' capabilities to reliably generate meta-actions in autonomous driving scenes. RAD leverages a retrieval-augmented generation (RAG) pipeline to dynamically improve decision accuracy through a three-stage process consisting of the embedding flow, retrieving flow, and generating flow. Additionally, we fine-tune VLMs on a specifically curated dataset derived from the NuScenes dataset to enhance their spatial perception and bird's-eye view image comprehension capabilities. Extensive experimental evaluations on the curated NuScenes-based dataset demonstrate that RAD outperforms baseline methods across key evaluation metrics, including match accuracy, and F1 score, and self-defined overall score, highlighting its effectiveness in improving meta-action decision-making for autonomous driving tasks.

自动驾驶视觉语言模型检索增强元动作决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。