arXiv:2501.17422cs.CVstat.AP2025-01

用统计模型预测图像凝视时长,生成真实凝视热图。

SIGN: A Statistically-Informed Gaze Network for Gaze Time Prediction

  • 结合统计建模与CNN/视觉变换器,预测整体凝视时间。
  • 在两个数据集上显著优于现有深度学习模型。
  • 可还原真实注视模式,适合视觉注意力研究者。

我们提出首个SIGN(Statistically-Informed Gaze Network)模型,用于预测图像上的总体凝视时长。通过构建基础统计模型,并实现基于卷积神经网络(CNN)和视觉变换器的深度学习架构,该模型能够预测整体凝视时间,并从中推导出各图像区域的凝视概率分布图,反映不同区域被注视的可能性。我们在AdGaze3500(包含广告图像的凝视时长数据集)和COCO-Search18(包含个体级注视点数据的搜索任务数据集)上评估SIGN性能。结果表明,SIGN在两个数据集上均显著优于当前最先进的深度学习基准方法;同时,在COCO-Search18中能生成与实际注视模式高度一致的凝视热图。这些结果证明SIGN首版已具备良好潜力,值得进一步开发。

原文摘要 · Abstract (English)

We propose a first version of SIGN, a Statistically-Informed Gaze Network, to predict aggregate gaze times on images. We develop a foundational statistical model for which we derive a deep learning implementation involving CNNs and Visual Transformers, which enables the prediction of overall gaze times. The model enables us to derive from the aggregate gaze times the underlying gaze pattern as a probability map over all regions in the image, where each region's probability represents the likelihood of being gazed at across all possible scan-paths. We test SIGN's performance on AdGaze3500, a dataset of images of ads with aggregate gaze times, and on COCO-Search18, a dataset with individual-level fixation patterns collected during search. We demonstrate that SIGN (1) improves gaze duration prediction significantly over state-of-the-art deep learning benchmarks on both datasets, and (2) can deliver plausible gaze patterns that correspond to empirical fixation patterns in COCO-Search18. These results suggest that the first version of SIGN holds promise for gaze-time predictions and deserves further development.

凝视预测视觉注意力概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。