arXiv:2608.07282cs.CLcs.CV2026-08

用现成的图文模型就能模拟人类看图说话时的注视行为。

Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders

论文配图:Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders
图 1 · 摘自论文原文
  • 用CLIP类双编码器加双模态归因方法预测注视点。
  • 无需微调或生成结构,就复现了经典英语视觉世界实验结果。
  • 适合对认知建模和多模态理解感兴趣的读者。

神经语言模型的进展也推动了计算心理语言学的发展,探讨这些模型是否能作为人类语言处理的潜在模型。然而,现有研究几乎都集中在书面或口语的单模态场景。相比之下,如视觉世界实验这类同时呈现视觉与语言输入的多模态范式却未受重视。本文提出一种新方法,通过结合CLIP类的简单多模态双编码器模型与双模态归因方法,预测视觉世界实验中的注视行为。我们证明该方法能稳健复现一项经典英语视觉世界研究的结果,展示了人类的预测性加工能力。令人惊讶的是,该方法既无生成架构,也无需针对此任务进行微调,即便如此仍表现良好。

原文摘要 · Abstract (English)

The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.

多模态认知建模注视预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。