arXiv:2411.13076cs.CV2024-11ICCV被引 7

通过提示增强提升多模态模型在自动驾驶中的视觉表征能力

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

  • 引入三类提示:关联、语义和问题提示,强化驾驶场景表征
  • 在有限数据下显著提升复杂交互与长尾场景的识别准确率
  • 适合自动驾驶多模态感知任务,尤其关注视觉-语言对齐

鉴于自动驾驶环境的动态性与严苛的安全要求,仅依赖CLIP的通用多模态大模型常难以准确表征驾驶特定场景,尤其是在复杂交互与长尾案例中表现不佳。为此,我们提出提示增强框架HoP,包含三项关键改进:关联提示通过强化标记间连接突出实例级结构;语义提示融入车辆间复杂互动、交通标志等驾驶相关高层信息;问题提示将视觉特征对齐至查询上下文,聚焦问题相关区域。这些提示通过提示融合模块融合,以有限领域数据捕捉驾驶相关表征,实现对驾驶场景的快速适应。大量实验表明,该框架在所有关键指标上均显著优于此前最优方法。

原文摘要 · Abstract (English)

In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hints of Prompt (HoP) framework, which introduces three key enhancements: Affinity hint to emphasize instance-level structure by strengthening token-wise connections, Semantic hint to incorporate high-level information relevant to driving-specific cases, such as complex interactions among vehicles and traffic signs, and Question hint to align visual features with the query context, focusing on question-relevant regions. These hints are fused through a Hint Fusion module, enriching visual representations by capturing driving-related representations with limited domain data, ensuring faster adaptation to driving scenarios. Extensive experiments confirm the effectiveness of the HoP framework, showing that it significantly outperforms previous state-of-the-art methods in all key metrics.

多模态自动驾驶视觉表征提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。