大模型无需注视训练也能自发关注任务目标,表现接近人类。
Emergent Goal-Directed Attention in Large Vision-Language Models
- 用自然场景测试大模型在搜索与自由观看下的注意力
- 目标匹配时模型注视点与人眼高度一致,即使无目标也成立
- 模型思考过程反映目标语义或视觉显著性,适合研究注意力机制
人类观察者会根据任务目标优先处理视觉信息。大多数自然场景注视建模依赖自由观看的注视数据训练,未明确目标导向注意是否能在无注视监督的系统中自发出现。我们测试了两个现成的视觉语言模型(VLMs):Qwen3-VL-32B-Thinking 和 Gemma-4-26B-A4B-it,基于 4,887 幅自然场景图像,在视觉搜索和自由观看指令下进行评估。将模型预测的注视点与人类在同一图像上的实际注视点对比。结果显示,当任务目标匹配时,模型注视点与人类更一致;在目标不存在的场景中,这种一致性依然存在,排除了单纯视觉锚定的解释,并在解码器层输出中可见。模型的思维轨迹在搜索任务中关联目标语义,在自由观看中关联视觉显著性。结果表明,通用型视觉语言模型可在无注视监督情况下生成符合人类行为的目标导向空间注意力,为理解目标导向注意提供新视角,并可作为跨任务预测人类注视的可扩展工具。
原文摘要 · Abstract (English)
Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free viewing, leaving open whether goal-directed attention can emerge in systems without gaze supervision. We tested two off-the-shelf vision-language models (VLMs), Qwen3-VL-32B-Thinking and Gemma-4-26B-A4B-it, on 4,887 naturalistic scenes under visual-search and free-viewing instructions. Model predictions were compared with human fixations on the same images under corresponding tasks. Both models aligned more closely with human fixations under matching goals than under mismatched goals. This crossover persisted in target-absent scenes, where alignment could not be explained by simple visual grounding, and appeared in decoder-layer readouts. Furthermore, model-thinking traces were grounded in target semantics during search and in visual prominence during free viewing. These findings show that general-purpose VLMs can generate human-aligned, goal-directed spatial priorities without gaze-specific training, informing theories of goal-directed attention and offering scalable tools for predicting where people look across tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。