arXiv:2511.06348cs.CVcs.AI2025-11被引 8

首个统一视觉与语言的多任务注视理解模型,提升注意力分析精度。

GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding

  • 融合RGB与深度图,通过文本提示实现多任务协同
  • 在GazeFollow和VideoAttentionTarget上达到当前最佳性能
  • 提出对象级注视检测指标,更精准评估目标识别效果

注视理解将人物检测、注视目标定位和兴趣对象识别整合到统一框架中,为视觉注意力与意图估计提供关键洞察。尽管已有研究建模视觉场景中的注视线索,但结合视觉与语言提示的统一系统仍需发展。本文提出GazeVLM,一种新型视觉-语言模型(VLM),用于图像中的多任务注视理解,涵盖人物检测、注视目标检测与注视对象识别。相较于其他基于Transformer的方法,GazeVLM是首个将VLM应用于这些联合任务的模型,支持任务选择性执行。通过融合视觉(RGB与深度)和文本模态,消融实验表明,在文本提示引导下,使用RGB图像与HHA编码深度图的组合表现最优。我们还引入对象级注视检测评估指标$AP_{ob}$。实验显示,GazeVLM在GazeFollow与VideoAttentionTarget数据集上显著提升,达到当前最优结果。

原文摘要 · Abstract (English)

Gaze understanding unifies the detection of people, their gaze targets, and objects of interest into a single framework, offering critical insight into visual attention and intent estimation. Although prior research has modelled gaze cues in visual scenes, a unified system is still needed for gaze understanding using both visual and language prompts. This paper introduces GazeVLM, a novel Vision-Language Model (VLM) for multi-task gaze understanding in images, addressing person detection, gaze target detection, and gaze object identification. While other transformer-based methods exist for gaze analysis, GazeVLM represents, to our knowledge, the first application of a VLM to these combined tasks, allowing for selective execution of each task. Through the integration of visual (RGB and depth) and textual modalities, our ablation study on visual input combinations revealed that a fusion of RGB images with HHA-encoded depth maps, guided by text prompts, yields superior performance. We also introduce an object-level gaze detection metric for gaze object identification ($AP_{ob}$). Through experiments, GazeVLM demonstrates significant improvements, notably achieving state-of-the-art evaluation scores on GazeFollow and VideoAttentionTarget datasets.

注视理解视觉语言模型多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。