VLMs常把头朝向误认作视线方向,影响人机交互理解。
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
- 用真实场景照片测试模型对视线目标的判断能力
- 模型准确率远低于人类,且依赖头朝向而非眼神特征
- 适合关注视觉语言模型偏差与非语言交流的研究者
视线是儿童和成人普遍使用的非语言沟通线索。视觉-语言模型(VLMs)能多好地推断视线目标?我们采集了1,360张真实场景照片,其中一人注视桌面上多个物体之一。关键在于,我们控制了观察者的头朝向:有时朝向注视目标,有时朝向干扰物,有时无约束。结果发现VLMs性能显著低于人类,排除了分辨率和物体命名等替代解释,并确定主因是模型将头朝向当作视线方向,而非基于眼睛外观判断。这种偏差可能源于数据而非架构,因为一个基于Transformer的视觉模型微调实验已验证此现象。未来研究应检验该发现是否普遍适用于各类深度学习方法,以及更优数据能否缓解所有架构的此类问题。明确原因有助于开发能更高效理解人类视线的交互技术。
原文摘要 · Abstract (English)
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation stimuli, we captured 1,360 real-world photos of scenes in which a person gazes at one of several objects on a table. Importantly, we also controlled the gazer's head orientation: sometimes it was directed toward the gaze target, sometimes toward a distractor object, and sometimes left unconstrained. We found a substantial performance gap between VLMs and humans, ruled out alternative explanations such as resolution and object-naming skills, and identified the main reason for the gap as VLMs inferring gaze direction using head orientation rather than eye appearance. Such a bias is likely due to data rather than architecture, as suggested by a proof-of-concept experiment finetuning a transformer-based vision model. Future work should investigate whether these findings hold broadly across various deep learning methods trained on existing data, and whether better data mitigates this problem for all architectures. Pinpointing the reason sets the stage for technologies that can interpret gaze targets to have more efficient interactions with humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。