arXiv:2510.19678cs.CVcs.AI2025-10被引 1

用视觉搜索测试大模型的感知能力,发现其像人一样有特征优先检测现象。

I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs

  • 借鉴认知心理学,用视觉搜索实验测试模型对颜色、大小等特征的感知机制。
  • 先进模型在单一特征搜索中表现出人类般的'突显效应',多特征搜索有容量限制。
  • 模型会利用自然场景光照先验,适合研究视觉理解机制的研究者参考。

多模态大语言模型(MLLM)在视觉-语言任务上表现优异,但其视觉处理过程不透明。现有黑箱评估仅关注任务准确率,难以揭示内在机制。受认知心理学启发,我们借鉴经典视觉搜索范式——最初用于研究人类感知——来检验MLLM是否具备“突显效应”,即显著视觉特征的检测不受干扰项数量影响。通过控制实验,针对颜色、大小和光照特征进行测试,发现先进MLLM在基于颜色或大小的分离式(单特征)搜索中表现出类人突显效应,且在组合式(多特征)搜索中存在容量限制。此外,还发现模型如人类一样将光照方向等自然场景先验融入物体表征。我们通过定向微调和机制可解释性分析进一步验证了这些发现。研究表明,视觉搜索可作为评估MLLM感知能力的认知基础诊断工具。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accuracy, but reveal little about underlying mechanisms. Drawing on cognitive psychology, we adapt classic visual search paradigms -- originally developed to study human perception -- to test whether MLLMs exhibit the ``pop-out'' effect, where salient visual features are detected independently of distractor set size. Using controlled experiments targeting colour, size and lighting features, we find that advanced MLLMs exhibit human-like pop-out effects in colour or size-based disjunctive (single feature) search, as well as capacity limits for conjunctive (multiple feature) search. We also find evidence to suggest that MLLMs, like humans, incorporate natural scene priors such as lighting direction into object representations. We reinforce our findings using targeted fine-tuning and mechanistic interpretability analyses. Our work shows how visual search can serve as a cognitively grounded diagnostic tool for evaluating perceptual capabilities in MLLMs.

多模态模型视觉搜索感知机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。