arXiv:2608.16514cs.CVcs.AI2026-08中稿 · ECCV

对比大模型与人类在视觉搜索中的注视模式,发现模型匹配结果但不复制过程。

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

论文配图:Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
图 1 · 摘自论文原文
  • 用人类匹配的视野逐点驱动模型,模拟真实搜索路径。
  • 模型检测目标接近人类上限,首跳即命中率更高。
  • 模型注视路径无时间序列性,非串行处理,与人类本质不同。

人类视觉搜索是串行的:中央凹必须落在候选物上才能确认,这些落点构成扫视路径。当多模态大语言模型(MLLMs)接收相同中央凹输入时,其搜索方式是否与人类一致,影响其作为人类视觉模型的可靠性及注意力对齐评分的有效性。我们对比了三种通用型MLLMs与人类眼动扫视路径在目标导向搜索任务(COCO-Search18)中的表现,通过逐次固定点驱动模型,评估三个维度:目标存在性判断、到达目标的效率以及注视过程本身。这三个维度可分离。在判断和目标获取方面,模型表现匹配或超越人类,对存在目标的检测接近满分,首次扫视命中率高于人类。然而注视过程并非人类式。在人类匹配条件下,所有三个模型均表现出一个特征:低熵、大振幅、自一致的扫视路径,自我一致性远超两人之间的一致性。这符合单次通过、非串行架构,而非视力限制所致。匹配视网膜输入能复现人类注视位置,但无法重现注视的时间展开过程;任何退化方案也无法恢复人类般的搜索行为与成功率。差异存在于过程轴上,而答案对齐与显著性指标并未测量此维度。因此,这些指标无法证明模型具有类人视觉;零样本模型适用于结果与空间问题,却不适用于时间与过程层面的问题。

原文摘要 · Abstract (English)

Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.

视觉搜索大模型注视分析人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。