arXiv:2606.05115cs.CVcs.AI2026-06

让AI像孩子一样在真实时间流中学习视觉与语言关联。

Continual Visual and Verbal Learning Through a Child's Egocentric Input

论文配图:Continual Visual and Verbal Learning Through a Child's Egocentric Input
图 1 · 摘自论文原文
  • 用单次时间顺序处理视频流,模拟儿童真实学习过程。
  • 在SAYCam数据集上表现超越流式学习基线,接近离线训练上限。
  • 双缓冲机制分别保存视觉与图文记忆,提升长期学习能力。

儿童从连续、有时间结构的自我中心体验中学习词汇含义。近期研究显示神经网络也能从儿童视角视频中学习词与对象的对应关系,但通常对打乱的数据循环数百轮训练,与儿童实际经验不符。本文提出BabyCL,一种持续多模态学习框架,仅以一次时间顺序遍历SAYCam数据集,结合流式视觉表示学习与图像-文本对比目标。BabyCL采用多阶段时间分段策略,并设计独立管理视觉与多模态历史的双重回放缓冲区,通过共享主干网络联合优化三个对比损失。在相同优化预算下,BabyCL在SAYCam Labeled-S 4AFC基准上优于流式学习基线,显著缩小与离线训练上限的差距。消融实验表明性能提升对在线时间分段窗口长度和回放缓冲区淘汰策略均具有鲁棒性。结果表明,在更贴近儿童真实体验的训练条件下,有意义的词-对象映射仍可有效形成。

原文摘要 · Abstract (English)

Children learn the meanings of words from a continuous, temporally structured stream of egocentric experience. Recent work shows that neural networks can also learn word-referent mappings from a child's egocentric video recordings, but they cycle through the shuffled data for hundreds of epochs, contrasting with how children actually encounter their environment. We introduce BabyCL, a continual multimodal learning framework that processes the SAYCam dataset in a single chronological pass, combining streaming visual representation learning with an image-text contrastive objective. BabyCL combines a multi-stage temporal segmentation of the stream with a dual replay buffer that independently manages visual and multimodal histories, and it is jointly trained with three contrastive losses on a shared backbone. Under a matched optimization budget, BabyCL outperforms streaming learning baselines on the SAYCam Labeled-S 4AFC benchmark, substantially narrowing the gap to an upper bound of offline training. Ablations show that the gains are robust to the length of the online temporal segmentation window and the eviction rule of the replay buffer. Together, these results show that meaningful word-referent mappings can emerge under training conditions much closer to a child's actual experience.

持续学习多模态儿童认知视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。