arXiv:2506.22494cs.ROcs.CV2025-06中稿 · IEEE/RSJ Internati…被引 3

用注意力引导生成自动驾驶复杂场景的可解释说明

DriveBLIP2: Attention-Guided Explanation Generation for Complex Driving Scenarios

  • 基于BLIP2-OPT架构,引入注意力图生成器聚焦关键驾驶目标
  • 在DRAMA数据集上,各评价指标显著优于基线模型
  • 适合关注自动驾驶可解释性与实时决策理解的研究者

本文提出一种新框架DriveBLIP2,基于BLIP2-OPT架构,用于生成复杂驾驶场景中准确且上下文相关的解释。现有视觉语言模型在一般任务中表现良好,但在多物体环境中的实时应用(如自动驾驶)中难以准确识别关键对象。为此,本文设计了注意力图生成器,突出关键视频帧中与驾驶决策相关的重要目标。通过引导模型关注这些关键区域,生成的注意力图有助于产生清晰、相关的解释,使驾驶员更易理解车辆在紧急情况下的决策过程。在DRAMA数据集上的评估显示,所提方法在BLEU、ROUGE、CIDEr和SPICE等指标上均显著优于基线模型,验证了定向注意力机制在提升实时自动驾驶可解释性方面的潜力。

原文摘要 · Abstract (English)

This paper introduces a new framework, DriveBLIP2, built upon the BLIP2-OPT architecture, to generate accurate and contextually relevant explanations for emerging driving scenarios. While existing vision-language models perform well in general tasks, they encounter difficulties in understanding complex, multi-object environments, particularly in real-time applications such as autonomous driving, where the rapid identification of key objects is crucial. To address this limitation, an Attention Map Generator is proposed to highlight significant objects relevant to driving decisions within critical video frames. By directing the model's focus to these key regions, the generated attention map helps produce clear and relevant explanations, enabling drivers to better understand the vehicle's decision-making process in critical situations. Evaluations on the DRAMA dataset reveal significant improvements in explanation quality, as indicated by higher BLEU, ROUGE, CIDEr, and SPICE scores compared to baseline models. These findings underscore the potential of targeted attention mechanisms in vision-language models for enhancing explainability in real-time autonomous driving.

自动驾驶可解释性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。