用推理令牌数模拟人类反应时间,检验视觉语言模型的搜索行为
Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

- 用推理令牌数量衡量模型搜索努力,类比人类反应时
- 前沿模型在复杂搜索中保持准确,中等模型则退化到随机水平
- 模型与人类在搜索模式上既相似又相异,揭示认知差异
视觉搜索是研究视觉注意的经典范式:反应时间随项目数量变化可区分并行‘突显’搜索与串行注意力需求搜索。本文探究视觉语言模型(VLMs)是否表现出类似行为特征。将四种经典范式——特征搜索与组合搜索、空间构型(T-vs-L)搜索、计数任务以及倾斜/垂直搜索不对称性——呈现给当前前沿与中等水平模型。由于单次模型调用无反应时间,本文以每轮尝试所消耗的推理(“思考”)令牌数作为模型内部的搜索努力代理,并与公开的人类基准数据(Wolfe et al., 2010)对比。结果显示,模型复现了多个类人特征:特征搜索的耗时恒定,而组合搜索代价随集合大小上升;前沿模型在高难度任务中保持准确,中等模型则退化至随机水平;分辨率控制实验表明,组合搜索代价源于真正的搜索过程而非小形状辨识困难。但模型也表现出显著差异:目标存在时的搜索努力斜率超过目标不存在时,逆转了人类规律;计数任务中模型仍准确,而人类会失准;具有自适应推理能力的模型完全拒绝在检测任务中多步推理,导致一种模型表现为努力梯度,另一种则为准确率断崖。本文认为,心理物理学范式作为行为探针,可低成本且精准评估机器视觉认知,而分歧点与一致点同样具有洞察价值。
原文摘要 · Abstract (English)
Visual search has been one of the most productive paradigms in the study of visual attention: the way reaction time scales with the number of items distinguishes parallel, "pop-out" search from serial, attention-demanding search. I ask whether vision-language models (VLMs) exhibit the same behavioral signatures. I adapt four classic paradigms: feature versus conjunction search, spatial-configuration (T-vs-L) search, enumeration, and the tilted/vertical search asymmetry; and present them to current frontier and mid-tier models. Because a single model call has no reaction time, I use the number of reasoning ("thinking") tokens a model spends per trial as a within-model analog of search effort, and I compare against a large public human benchmark (Wolfe et al., 2010). The models reproduce several human signatures: feature search costs flat effort while conjunction effort climbs with set size; frontier models hold accuracy where mid-tier models collapse to chance; and a resolution control shows the conjunction cost is genuine search rather than difficulty resolving small shapes. They also diverge from humans in informative ways. The target-present effort slope exceeds the target-absent slope, reversing the human ordering; enumeration remains accurate where humans would lose count; and a reasoning model with adaptive deliberation declines to deliberate on detection tasks altogether, so that a single search expresses itself as an effort gradient in one model and as an accuracy cliff in another. I argue that psychophysical paradigms, applied behaviorally, are a sharp and inexpensive probe of machine visual cognition, and that the points of divergence are as informative as the points of agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。