arXiv:2511.11725cs.CVcs.AI2025-11

用人类视觉盲区设计新掩码,提升儿童语言学习模型效果

Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video

  • 基于人眼盲区设计掩码策略,模拟大脑补全视觉空缺机制
  • 在跨情境和时序扩展场景中,效果不弱于传统随机掩码
  • 适合研究儿童语言发展与生物启发视觉模型的学者

婴儿通常在6至9个月大时开始学习首个词汇,将口语表达与视觉参照物建立联系。由于缺乏先验知识,一个首次出现的词可能指向环境中任意物体、其组成部分或属性。本文利用一名儿童的纵向、第一视角且生态有效的视频数据,提出一种自监督且符合生物学原理的强视觉表征学习方法。基于掩码自编码器的视觉主干网络引入人眼盲区知识,设计新型掩码策略,模拟大脑填补视野空白的方式。该方法相较传统随机掩码更具生物学合理性。预训练编码器被用于基于对比学习的视频-文本模型,实现词与视觉参照物的映射。大量实验表明,该生物合理掩码策略在跨情境及时间延展性任务中,至少与随机掩码同样有效。

原文摘要 · Abstract (English)

Typically, children start to learn their first words between 6 and 9 months, linking spoken utterances to their visual referents. Without prior knowledge, a word encountered for the first time can be interpreted in countless ways; it might refer to any of the objects in the environment, their components, or attributes. Using longitudinal, egocentric, and ecologically valid data from the experience of one child, in this work, we propose a self-supervised and biologically plausible strategy to learn strong visual representations. Our masked autoencoder-based visual backbone incorporates knowledge about the blind spot in human eyes to define a novel masking strategy. This mask and reconstruct approach attempts to mimic the way the human brain fills the gaps in the eyes' field of view. This represents a significant shift from standard random masking strategies, which are difficult to justify from a biological perspective. The pretrained encoder is utilized in a contrastive learning-based video-text model capable of acquiring word-referent mappings. Extensive evaluation suggests that the proposed biologically plausible masking strategy is at least as effective as random masking for learning word-referent mappings from cross-situational and temporally extended episodes.

视觉建模语言学习生物启发自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。