arXiv:2505.12660cs.CV2025-05被引 1

用视觉聚焦机制和视觉语言模型预测人看图理解时间。

Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps

  • 结合人眼聚焦特性和视觉语言模型生成场景理解地图
  • 模型相关性达r=0.47(反应时间)和r=0.51(眼跳次数)
  • 适合研究认知心理学与视觉感知的学者

尽管已有模型可预测目标搜索或视觉辨别任务中的人类反应时间,但针对场景理解时间的图像可计算预测器仍属空白。近期视觉语言模型(VLMs)能为任意图像生成场景描述,结合语言描述的量化评估指标,为建模人类场景理解提供了新契机。我们假设:人类场景理解的主要瓶颈及反应时间差异的核心来源,是人眼聚焦特性与图像中任务相关信息的空间分布之间的交互作用。基于此,我们提出一种新型图像可计算模型,将聚焦视觉与VLM结合,生成随注视位置变化的场景理解空间映射(聚焦场景理解图,F-SUM),并引入综合得分。该得分与17名被试的平均反应时间(r=0.47)和277个场景中的眼跳次数(r=0.51)显著相关,且与16名被试在限时呈现下的描述准确率(r=-0.56)也显著相关。这些相关性远超传统图像度量如杂乱度、视觉复杂度和基于语言熵的场景模糊度。本工作引入了一种新的图像可计算指标,用于预测场景理解反应时间,并揭示了聚焦视觉处理在理解难度塑造中的关键作用。

原文摘要 · Abstract (English)

Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene understanding time remains an open challenge. Recent advances in vision-language models (VLMs), which can generate scene descriptions for arbitrary images, combined with the availability of quantitative metrics for comparing linguistic descriptions, offer a new opportunity to model human scene understanding. We hypothesize that the primary bottleneck in human scene understanding and the driving source of variability in response times across scenes is the interaction between the foveated nature of the human visual system and the spatial distribution of task-relevant visual information within an image. Based on this assumption, we propose a novel image-computable model that integrates foveated vision with VLMs to produce a spatially resolved map of scene understanding as a function of fixation location (Foveated Scene Understanding Map, or F-SUM), along with an aggregate F-SUM score. This metric correlates with average (N=17) human RTs (r=0.47) and number of saccades (r=0.51) required to comprehend a scene (across 277 scenes). The F-SUM score also correlates with average (N=16) human description accuracy (r=-0.56) in time-limited presentations. These correlations significantly exceed those of standard image-based metrics such as clutter, visual complexity, and scene ambiguity based on language entropy. Together, our work introduces a new image-computable metric for predicting human response times in scene understanding and demonstrates the importance of foveated visual processing in shaping comprehension difficulty.

认知模型视觉语言反应时间眼动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。