首个视觉语言模型眼神理解基准,验证大模型能否从通用训练中学会看人眼神。
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- 构建48.9万条问答数据,将眼神理解统一为VQA任务
- 大模型在无特定训练下难以准确识别眼神指向与位置
- 针对性多任务训练可显著提升眼神理解能力,适合智能交互研究者
人类眼神是理解注意力、意图与社交互动的关键线索,但当前视觉语言模型(VLMs)对眼神的理解仍基本空白。尽管近期的VLMs在多种视觉任务上表现强劲,却缺乏系统评估或训练其进行眼神解读的基准。为此,我们提出VL4Gaze,首个大规模基准,用于探究、评估并激发VLMs在眼神理解方面的潜力。该数据集包含124,000张图像上的489,000条自动生成的问答对,通过四项互补任务将眼神理解建模为统一的视觉问答问题:(1) 眼神注视物体描述,(2) 眼神方向描述,(3) 眼神定位,(4) 模糊问题识别。我们在上下文学习与微调设置下全面评估商业及开源VLMs。结果显示,即使大规模模型在无特定监督时也难以可靠推断眼神语义与空间定位。相比之下,在VL4Gaze上训练后,所有任务均获得显著且一致的性能提升,凸显了针对多任务监督在发展眼神理解能力中的关键作用。我们将公开数据集与代码,以支持此方向的进一步研究与开发。
原文摘要 · Abstract (English)
Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current vision-language models (VLMs). While recent VLMs achieve strong scene-level reasoning across a range of visual tasks, there exists no benchmark that systematically evaluates or trains them for gaze interpretation, leaving open the question of whether gaze understanding can emerge from general-purpose vision-language pre-training. To address this gap, we introduce VL4Gaze, the first large-scale benchmark designed to investigate, evaluate, and unlock the potential of VLMs for gaze understanding. VL4Gaze contains 489K automatically generated question-answer pairs across 124K images and formulates gaze understanding as a unified VQA problem through four complementary tasks: (1) gaze object description, (2) gaze direction description, (3) gaze point location, and (4) ambiguous question recognition. We comprehensively evaluate both commercial and open-source VLMs under in-context learning and fine-tuning settings. The results show that even large-scale VLMs struggle to reliably infer gaze semantics and spatial localization without task-specific supervision. In contrast, training on VL4Gaze brings substantial and consistent improvements across all tasks, highlighting the importance of targeted multi-task supervision for developing gaze understanding capabilities in VLMs. We will release the dataset and code to support further research and development in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。