评测视觉语言模型对人类目光追踪与社交注视的理解能力
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

- 构建双维度评估框架,涵盖注视跟随与社交注视预测任务
- 现有VLM在目光定位上表现不佳,远低于纯视觉模型
- 提示工程与微调均能提升性能,但仍有显著差距
视觉语言模型(VLMs)已发展为具备强大零样本泛化能力的通用多模态推理器。本文提出EyeVLM,系统评估VLM在目光理解方面的表现,涵盖两个互补任务:一是注视跟随(预测人看的位置),侧重几何与视觉感知;二是社交注视预测(推断多人互动中的相互凝视与共同关注),依赖社会关系推理。评估覆盖多种开源与闭源SOTA VLM,在零样本和微调两种设置下比较不同提示策略、模型规模与数据规模的影响。基于现有眼动数据集,与纯视觉模型进行系统对比。结果表明,当前VLM在精确目光理解方面仍严重不足,尽管标准训练可缩小差距,但需大幅提升性能。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human behavior understanding that requires reasoning about the physical scene as well as the activity, interactions, and social context. However, the extent to which VLMs can reliably understand human gaze and related attentional behaviors remains largely unexplored. In this work, we present EyeVLM, a systematic evaluation framework for gaze understanding in VLMs across two complementary dimensions: tasks and models. To assess gaze understanding capabilities, we focus on two core tasks. The first, gaze following, i.e., predicting the 2D location where a person is looking, has a geometric and visual processing focus, requiring a precise understanding of the human face, attention direction, 3D scene structure, and spatial grounding of attended targets. The second, social gaze prediction, requires social and relational reasoning over multi-person interactions (e.g., mutual gaze and shared attention), and may benefit more from the LLM semantic reasoning capabilities within VLMs. Regarding models, EyeVLM evaluates these tasks in two ways: a zero-shot setting with a diverse set of state-of-the-art open- and closed-source VLMs, exploring different prompting strategies; and a fine-tuning approach based on task-specific QA pairs, studying the impact of model scale and data scale. As benchmarks, we rely on existing gaze understanding datasets and perform a systematic comparison with state-of-the-art purely visual models. Overall, our results show that current VLMs lack precise gaze understanding capabilities. While standard training helps reduce the gap with visual models, significant improvements are still needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。