arXiv:2502.18042cs.CVcs.AI2025-02被引 27

用视觉语言模型增强自动驾驶注意力机制,提升复杂场景下的感知与决策能力。

VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion

  • 引入文本特征融合到鸟瞰图特征中,用语义信息指导模型学习人类注意力模式。
  • 在nuScenes数据集上,感知、预测和规划任务性能显著优于基线模型。
  • 提出可学习加权融合策略,动态平衡视觉与文本模态贡献,缓解模态重要性失衡。

人类驾驶员能通过丰富的注意力语义应对复杂场景,但现有自动驾驶系统在将2D观测转换为3D空间时往往丢失关键语义信息,限制其在动态复杂环境中的部署。为此,我们提出VLM-E2E框架,利用视觉语言模型(VLMs)提供注意力提示以增强训练。该方法将文本表示融入鸟瞰图(BEV)特征,实现语义监督,使模型学习更丰富的特征表示,显式捕捉驾驶者的注意力语义。通过聚焦注意力语义,VLM-E2E更贴近人类驾驶行为,有助于应对复杂动态环境。此外,我们设计了BEV-Text可学习加权融合策略,解决多模态信息融合中的模态重要性失衡问题,动态平衡视觉与文本特征的贡献,确保互补信息被有效利用。在nuScenes数据集上的实验表明,相比基线端到端模型,VLM-E2E在感知、预测和规划任务上均有显著提升,验证了注意力增强的BEV表示在提升自动驾驶准确性与可靠性方面的有效性。

原文摘要 · Abstract (English)

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D observations into 3D space. In this sense, it hinders their effective deployment in dynamic and complex environments. Leveraging the superior scene understanding and reasoning abilities of Vision-Language Models (VLMs), we propose VLM-E2E, a novel framework that uses the VLMs to enhance training by providing attentional cues. Our method integrates textual representations into Bird's-Eye-View (BEV) features for semantic supervision, which enables the model to learn richer feature representations that explicitly capture the driver's attentional semantics. By focusing on attentional semantics, VLM-E2E better aligns with human-like driving behavior, which is critical for navigating dynamic and complex environments. Furthermore, we introduce a BEV-Text learnable weighted fusion strategy to address the issue of modality importance imbalance in fusing multimodal information. This approach dynamically balances the contributions of BEV and text features, ensuring that the complementary information from visual and textual modalities is effectively utilized. By explicitly addressing the imbalance in multimodal fusion, our method facilitates a more holistic and robust representation of driving environments. We evaluate VLM-E2E on the nuScenes dataset and achieve significant improvements in perception, prediction, and planning over the baseline end-to-end model, showcasing the effectiveness of our attention-enhanced BEV representation in enabling more accurate and reliable autonomous driving tasks.

自动驾驶多模态融合视觉语言模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。