模拟人眼聚焦机制的模型在理解场景时自然出现人类注视模式。
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

- 用模拟中央视觉的模型优化场景理解,自发产生人类注视模式。
- 该模型预测人类注视点准确率高于其他训练目标或视觉配置。
- 适合关注视觉认知机制与生物启发模型的研究者。
当人类在无特定任务下自由观察场景时,会先看向中心,随后聚焦于人物、文字、被凝视或抓握的物体以及语义有意义区域。这些典型注视模式反映什么,是否优化了某种感知任务尚不清楚。我们发现,一个具有模拟中央视觉的计算代理,在训练以优化场景理解时,会自发产生类似人类的注视模式。相比之下,以搜索或分类任务训练的代理,或配备优于/劣于人类周边视觉的代理,预测人类注视点的准确性更低。因此,人类自由注视模式可能是优化场景理解在中央视觉生物约束下的功能性副产物。
原文摘要 · Abstract (English)
When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on people, text, objects being gazed at or grasped, and semantically meaningful regions. What these signature fixation patterns reflect and whether they optimize an underlying perceptual task remain unknown. We show that a computational agent with simulated foveation, trained to optimize scene comprehension, exhibits emergent human fixation signature patterns. In contrast, versions of the agent trained to search or classify scenes, or equipped with peripheral vision that was better or worse than human vision, predicted human fixation patterns less accurately. Thus, human free-viewing fixation patterns may emerge as a functional byproduct of optimizing scene comprehension under the biological constraints of foveated vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。