arXiv:2604.20191cs.CVcs.AI2026-04

提出双分支模型,实现文本引导的精细驾驶注意力预测。

From Scene to Object: Text-Guided Dual-Gaze Prediction

论文配图:From Scene to Object: Text-Guided Dual-Gaze Prediction
图 1 · 摘自论文原文
  • 用大语言模型与SAM3构建细粒度注视数据集。
  • 在安全关键场景下相似度提升17.8%。
  • 生成的注意力图88.22%被人类认为真实可信。

可解释的驾驶员注意力预测对类人自动驾驶至关重要。然而现有数据集仅提供场景级全局注视,缺乏细粒度物体级标注,难以支持文本-语义建模。尽管视觉-语言模型(VLMs)具备语义推理潜力,但数据限制导致严重文本-视觉脱节和视觉偏见幻觉。为此,本文提出新型双分支注视预测框架,建立从数据构建到模型架构的完整范式。首先,构建G-W3DA物体级驾驶员注意力数据集,通过多模态大语言模型与Segment Anything Model 3(SAM3)结合,在严格交叉验证下将宏观热图解耦为物体级掩码,从根本上消除标注幻觉。基于此高质量数据,提出DualGaze-VLM架构,通过条件感知SE门控动态调制视觉特征,实现意图驱动的空间精确定位。在W3DA基准上的大量实验表明,DualGaze-VLM在空间对齐指标上持续优于现有SOTA模型,尤其在安全关键场景下相似度(SIM)最高提升17.8%。此外,视觉图灵测试显示,88.22%的人类评估者认为其生成的注意力热图具有真实性,证明其能生成合理的认知先验。

原文摘要 · Abstract (English)

Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaze rather than fine-grained object-level annotations, inherently failing to support text-grounded cognitive modeling. Consequently, while Vision-Language Models (VLMs) hold great potential for semantic reasoning, this critical data limitations leads to severe text-vision decoupling and visual-bias hallucinations. To break this bottleneck and achieve precise object-level attention prediction, this paper proposes a novel dual-branch gaze prediction framework, establishing a complete paradigm from data construction to model architecture. First, we construct G-W3DA, a object-level driver attention dataset. By integrating a multimodal large language model with the Segment Anything Model 3 (SAM3), we decouple macroscopic heatmaps into object-level masks under rigorous cross-validation, fundamentally eliminating annotation hallucinations. Building upon this high-quality data foundation, we propose the DualGaze-VLM architecture. This architecture extracts the hidden states of semantic queries and dynamically modulates visual features via a Condition-Aware SE-Gate, achieving intent-driven precise spatial anchoring. Extensive experiments on the W3DA benchmark demonstrate that DualGaze-VLM consistently surpasses existing state-of-the-art (SOTA) models in spatial alignment metrics, notably achieving up to a 17.8% improvement in Similarity (SIM) under safety-critical scenarios. Furthermore, a visual Turing test reveals that the attention heatmaps generated by DualGaze-VLM are perceived as authentic by 88.22% of human evaluators, proving its capability to generate rational cognitive priors.

注意力预测视觉语言模型自动驾驶数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。