arXiv:2501.02966cs.CVcs.LG2025-01被引 1

模拟人眼注视区域训练,提升视觉模型对物体的感知能力

Human Gaze Boosts Object-Centered Representation Learning

  • 用注视预测生成视觉中心区域,聚焦关键信息
  • 在Ego4D数据集上,中心视野训练使物体表征更优
  • 适合研究生物启发视觉学习的科研人员

近期基于人类视角视觉输入的自监督学习(SSL)模型在图像识别任务中表现远低于人类。这些模型使用头戴摄像头采集的原始、均匀视觉输入,而人类的视网膜与视觉皮层会放大注视点附近的中央视觉信息,这种选择性放大可能有助于形成以物体为中心的视觉表征。本文探究聚焦中央视觉是否能提升以我视角的视觉学习效果。我们利用大规模Ego4D数据集模拟5个月的视角经验,并通过人类注视预测模型生成注视位置。为体现人类中央视觉的重要性,我们裁剪注视点附近的视觉区域。最后,在这些处理后的输入上训练时间型自监督模型。实验表明,聚焦中央视觉可生成更优的物体中心表征。分析显示,该模型利用注视运动的时间动态构建了更强的视觉表示。本工作推动了类生物视觉表征学习的发展。

原文摘要 · Abstract (English)

Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models train on raw, uniform visual inputs collected from head-mounted cameras. This is different from humans, as the anatomical structure of the retina and visual cortex relatively amplifies the central visual information, i.e. around humans' gaze location. This selective amplification in humans likely aids in forming object-centered visual representations. Here, we investigate whether focusing on central visual information boosts egocentric visual object learning. We simulate 5-months of egocentric visual experience using the large-scale Ego4D dataset and generate gaze locations with a human gaze prediction model. To account for the importance of central vision in humans, we crop the visual area around the gaze location. Finally, we train a time-based SSL model on these modified inputs. Our experiments demonstrate that focusing on central vision leads to better object-centered representations. Our analysis shows that the SSL model leverages the temporal dynamics of the gaze movements to build stronger visual representations. Overall, our work marks a significant step toward bio-inspired learning of visual representations.

视觉表征自监督学习注视预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。