arXiv:2509.15751cs.CV2025-09中稿 · IEEE ICDL 2025被引 1

模拟人眼中央高分辨、周边低分辨的视觉特性,提升自监督模型对物体表征的学习效果。

Simulated Cortical Magnification Supports Self-Supervised Object Learning

  • 在视频输入中加入视网膜中心高分辨率、边缘低分辨率的模拟
  • 模型在真实人类视角视频上训练后,物体表征质量显著提升
  • 适合关注生物启发式视觉学习与自监督模型改进的研究者

近期自监督学习模型通过模拟幼儿的视觉经验来构建语义物体表征,但忽略了人眼视觉中心高分辨、周边低分辨的特点。本文研究这一变分辨率特性在物体表征发展中的作用。利用两个记录人类与物体交互过程的自然视角视频数据集,应用人眼视网膜聚焦与皮层放大模型对输入进行处理,使视觉内容向周边逐渐模糊。基于这些修改后的序列,训练两种受生物学启发的自监督模型,采用基于时间的学习目标。结果表明,建模人眼的这种变分辨率特性可提升所学物体表征的质量。分析显示,这一提升源于物体在中心区域显得更大,以及更优的中心与外围信息权衡。该工作推动了人类视觉表征学习模型在真实性和性能上的进步。

原文摘要 · Abstract (English)

Recent self-supervised learning models simulate the development of semantic object representations by training on visual experience similar to that of toddlers. However, these models ignore the foveated nature of human vision with high/low resolution in the center/periphery of the visual field. Here, we investigate the role of this varying resolution in the development of object representations. We leverage two datasets of egocentric videos that capture the visual experience of humans during interactions with objects. We apply models of human foveation and cortical magnification to modify these inputs, such that the visual content becomes less distinct towards the periphery. The resulting sequences are used to train two bio-inspired self-supervised learning models that implement a time-based learning objective. Our results show that modeling aspects of foveated vision improves the quality of the learned object representations in this setting. Our analysis suggests that this improvement comes from making objects appear bigger and inducing a better trade-off between central and peripheral visual information. Overall, this work takes a step towards making models of humans' learning of visual representations more realistic and performant.

自监督学习生物启发视觉表征认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。