用婴儿发育过程启发视觉语言模型预训练,提升样本效率。
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- 基于婴儿成长轨迹构建多模态数据集,模拟真实早期认知体验。
- 小型模型在10项发育任务中表现超越GPT-4o部分能力。
- 提供可评估认知能力的基准工具箱,适合发展性AI研究者。
早期儿童的发展轨迹为视觉基础模型的高效样本预训练提供了自然目标。我们提出BabyVLM-V2,一个基于婴儿认知发展的视觉语言建模框架,相较BabyVLM-V1在纵向多维度预训练数据集、通用模型结构以及核心的DevCV工具箱方面均有显著提升。预训练数据集覆盖广泛且人工标注少,包含视频-话语、图像-话语及多轮对话数据,忠实反映婴儿认知经验。DevCV工具箱将最新发布的NIH Baby Toolbox中的全部视觉相关评估指标转化为一套包含10个跨模态任务的基准测试,涵盖空间推理、记忆与词汇理解等早期儿童能力。实验表明,从零开始预训练的小型模型在该基准上表现优异,部分任务性能超过GPT-4o。我们希望这一系统化、统一的BabyVLM-V2框架能加速发展性合理视觉基础模型的研究。
原文摘要 · Abstract (English)
Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal, multifaceted pretraining set, a versatile model, and, most importantly, DevCV Toolbox for cognitive evaluation. The pretraining set maximizes coverage while minimizing curation of a longitudinal, infant-centric audiovisual corpus, yielding video-utterance, image-utterance, and multi-turn conversational data that mirror infant experiences. DevCV Toolbox adapts all vision-related measures of the recently released NIH Baby Toolbox into a benchmark suite of ten multimodal tasks, covering spatial reasoning, memory, and vocabulary understanding aligned with early children's capabilities. Experimental results show that a compact model pretrained from scratch can achieve competitive performance on DevCV Toolbox, outperforming GPT-4o on some tasks. We hope the principled, unified BabyVLM-V2 framework will accelerate research in developmentally plausible pretraining of vision foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。