arXiv:2510.09299cs.CVeess.IV2025-10

人类视觉扫描遵循类似动物觅食的最优搜索模式。

Foraging with the Eyes: Dynamics in Human Visual Gaze and Deep Predictive Modeling

  • 发现人眼注视轨迹符合莱维飞行规律,具高效搜索特性。
  • 40人看50图,超400万注视点数据验证该规律存在。
  • 仅凭图像就能用CNN预测注视热点,模型可复现关键行为。

动物常通过莱维飞行(重尾步长的随机轨迹)在稀疏资源环境中觅食。本研究发现,人类在浏览图像时的眼动轨迹也呈现类似动态。尽管传统模型强调图像显著性,但眼动的时空统计特性仍待深入探索。本研究开展大规模实验,40名参与者在自由条件下观看50张多样图像,使用高速眼动仪记录超过400万次注视点。数据分析表明,人眼注视轨迹同样遵循类莱维飞行模式,暗示人类视觉以最优效率搜寻信息。此外,我们训练了一个卷积神经网络(CNN),仅基于图像输入即可预测注视热图,模型在新图像上准确复现了显著注视区域,证明眼动关键特征可从视觉结构中学习。研究结果为视觉探索提供了新的统计规律支持,并为生成式与预测性眼动建模开辟了新路径。

原文摘要 · Abstract (English)

Animals often forage via Levy walks stochastic trajectories with heavy tailed step lengths optimized for sparse resource environments. We show that human visual gaze follows similar dynamics when scanning images. While traditional models emphasize image based saliency, the underlying spatiotemporal statistics of eye movements remain underexplored. Understanding these dynamics has broad applications in attention modeling and vision-based interfaces. In this study, we conducted a large scale human subject experiment involving 40 participants viewing 50 diverse images under unconstrained conditions, recording over 4 million gaze points using a high speed eye tracker. Analysis of these data shows that the gaze trajectory of the human eye also follows a Levy walk akin to animal foraging. This suggests that the human eye forages for visual information in an optimally efficient manner. Further, we trained a convolutional neural network (CNN) to predict fixation heatmaps from image input alone. The model accurately reproduced salient fixation regions across novel images, demonstrating that key components of gaze behavior are learnable from visual structure alone. Our findings present new evidence that human visual exploration obeys statistical laws analogous to natural foraging and open avenues for modeling gaze through generative and predictive frameworks.

视觉注意力眼动追踪莱维飞行深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。