arXiv:2602.04462cs.CV2026-02

聚焦中心视觉与时间缓慢性,提升模型对物体语义的感知能力

Temporal Slowness in Central Vision Drives Semantic Object Learning

  • 基于预测眼动轨迹提取中心视觉区域,结合时间对比自监督学习
  • 在Ego4D数据集上训练后,对象特征提取能力显著增强
  • 适合关注视觉认知机制与自监督学习的研究者

人类在极少监督下从第一人称视觉流中习得语义物体表征,但其内在机制尚不明确。视觉系统仅以高分辨率处理视野中心,且对时间上接近的视觉输入形成相似表征,强调注视点附近缓慢变化的信息。本研究通过Ego4D数据集与先进眼动预测模型,模拟五个月类人视觉经验,提取预测注视点周围的图像块,训练时间对比自监督学习模型。结果表明,在中心视觉经验中利用时间缓慢性可改善对物体语义多方面的编码。具体而言,聚焦中心视觉强化了前景物体特征提取,而结合眼动的时间缓慢性则有助于捕捉更广泛的物体语义信息。该研究为人类如何从自然视觉体验中构建语义表征提供了新见解。代码将在录用后公开,地址:https://github.com/t9s9/central-vision-ssl。

原文摘要 · Abstract (English)

Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system only processes the center of its field of view with high resolution and it learns similar representations for visual inputs occurring close in time. This emphasizes slowly changing information around gaze locations. This study investigates the role of central vision and slowness learning in the formation of semantic object representations from human-like visual experience. We simulate five months of human-like visual experience using the Ego4D dataset and a state-of-the-art gaze prediction model. We extract image crops around predicted gaze locations to train a time-contrastive Self-Supervised Learning model. Our results show that exploiting temporal slowness when learning from central visual field experience improves the encoding of different facets of object semantics. Specifically, focusing on central vision strengthens the extraction of foreground object features, while considering temporal slowness, especially in conjunction with eye movements, allows the model to encode broader semantic information about objects. These findings provide new insights into the mechanisms by which humans may develop semantic object representations from natural visual experience. Our code will be made public upon acceptance. Code is available at https://github.com/t9s9/central-vision-ssl.

视觉认知自监督学习眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。