arXiv:2411.10639cs.CVcs.AI2024-11被引 7

让鸟瞰图感知与语言描述相互对齐,提升自动驾驶环境理解能力

MTA: Multimodal Task Alignment for BEV Perception and Captioning

  • 通过上下文对齐机制,让鸟瞰图与语言表示匹配
  • 在nuScenes和TOD3Cap上感知与描述任务分别提升10.7%和9.2%
  • 无需额外计算开销,可无缝融入现有模型

基于鸟瞰图(BEV)的3D感知在自动驾驶中至关重要。大型语言模型的兴起推动了基于BEV的场景描述研究,以理解周围物体的行为。然而,现有方法将感知与描述视为独立任务,仅关注单一任务性能,忽视了多模态对齐的潜力。为此,本文提出MTA——一种新型多模态任务对齐框架,同时提升BEV感知与描述能力。MTA包含两个核心组件:(1) 鸟瞰图-语言对齐(BLA),通过上下文学习使BEV场景表征与真实语言表征对齐;(2) 检测-描述对齐(DCA),通过跨模态提示机制对齐检测与描述输出。MTA在训练阶段无缝集成至先进基线,运行时无额外计算开销。在nuScenes和TOD3Cap数据集上的大量实验表明,MTA在两项任务上均显著优于现有最先进方法,罕见场景下感知性能提升10.7%,描述性能提升9.2%。结果验证了统一对齐在协调BEV感知与描述中的有效性。

原文摘要 · Abstract (English)

Bird's eye view (BEV)-based 3D perception plays a crucial role in autonomous driving applications. The rise of large language models has spurred interest in BEV-based captioning to understand object behavior in the surrounding environment. However, existing approaches treat perception and captioning as separate tasks, focusing on the performance of only one task and overlooking the potential benefits of multimodal alignment. To bridge this gap between modalities, we introduce MTA, a novel multimodal task alignment framework that boosts both BEV perception and captioning. MTA consists of two key components: (1) BEV-Language Alignment (BLA), a contextual learning mechanism that aligns the BEV scene representations with ground-truth language representations, and (2) Detection-Captioning Alignment (DCA), a cross-modal prompting mechanism that aligns detection and captioning outputs. MTA seamlessly integrates into state-of-the-art baselines during training, adding no extra computational complexity at runtime. Extensive experiments on the nuScenes and TOD3Cap datasets show that MTA significantly outperforms state-of-the-art baselines in both tasks, achieving a 10.7% improvement in challenging rare perception scenarios and a 9.2% improvement in captioning. These results underscore the effectiveness of unified alignment in reconciling BEV-based perception and captioning.

BEV感知多模态对齐自动驾驶语言描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。