arXiv:2605.14990cs.CV2026-05

分析婴幼儿视觉经验发现:孩子看物体很杂乱,却仍能精准分类。

Characterizing the visual representation of objects from the child's view

论文配图:Characterizing the visual representation of objects from the child's view
图 1 · 摘自论文原文
  • 用300万帧第一视角视频提取儿童常见物体,揭示其视觉输入高度不均衡。
  • 多数物体仅偶尔出现,且常以异常角度、遮挡或图画形式呈现。
  • 尽管输入混乱,孩子仍能按类别(如动物、食物)强关联,适合研究鲁棒学习模型。

儿童在生命前几岁通过日常经验习得物体类别表征。这些学习过程的输入长什么样?我们基于BabyView数据集(N=31名参与者,868小时,年龄5-36个月)的第一人称视频,利用监督物体检测模型从超过300万帧中提取常见物体类别。发现儿童的物体暴露高度偏斜:少数类别(如杯子、椅子)占据主导,多数类别出现频率极低,与以往小范围环境研究结果一致。物体实例高度可变:儿童常以非常规角度、在高度杂乱场景中、部分遮挡下观察物体;许多类别(尤其是动物)最常以图像形式被看到。令人意外的是,尽管输入高度变异,检测到的类别(如长颈鹿、苹果)在超类别(如动物、食物)内的分组强度反而高于来自标准照片的分组。这一模式在自监督视觉与多模态模型的高维嵌入中同样成立;在个体儿童密集采样的数据中也复现。理解视觉类别学习的鲁棒性与效率,需要开发能利用强超类别结构、并从非标准、稀疏、可变实例中学习的模型。

原文摘要 · Abstract (English)

Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process look like? We analyzed first-person videos of young children's visual experience at home from the BabyView dataset ($N$ = 31 participants, 868 hours, ages 5--36 months), using a supervised object detection model to extract common object categories from more than 3 million frames. We found that children's object category exposure was highly skewed: a few categories (e.g., cups, chairs) dominated children's visual experiences while most categories appeared rarely, replicating previous findings from a more restricted set of contexts. Category exemplars were highly variable: children encountered objects from unusual angles, in highly cluttered scenes, and partially occluded views; many categories (especially animals) were most frequently viewed as depictions. Surprisingly, despite this variability, detected categories (e.g., giraffes, apples) showed stronger groupings within superordinate categories (e.g., animals, food) relative to groupings derived from canonical photographs of these categories. We found this same pattern when using high-dimensional embeddings from both self-supervised visual and multimodal models; this effect was also recapitulated in densely sampled data from individual children. Understanding the robustness and efficiency of visual category learning will require the development of models that can exploit strong superordinate structure and learn from non-canonical, sparse, and variable exemplars.

视觉认知儿童发展物体识别非标准数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。