arXiv:2504.09426cs.CVcs.AI2025-04ICCV被引 8

模仿婴儿学习方式,用少量数据训练出更高效视觉语言模型

BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning

  • 用儿童导向变换生成合成数据,模拟婴儿学习环境
  • 在小规模数据上训练的模型性能超越仅用SAYCam数据的模型
  • 适合研究数据效率与儿童认知启发的AI学习者

人类婴儿能从极少输入中快速发展出视觉推理能力,提示以发育过程为灵感的预训练可显著提升视觉语言模型(VLMs)的数据效率。尽管近期已有基于婴儿启发数据集如SAYCam的研究,但现有评估基准仍存在偏差——或过于简单、范围狭窄,或专为大规模预训练模型设计。此外,仅使用婴儿数据忽略了婴儿自然学习中丰富的多样化输入。为此,我们提出BabyVLM,包含全面的领域内评估基准和通过现有数据的儿童导向变换生成的合成训练数据集。实验表明,使用该合成数据集训练的VLM在BabyVLM任务上的表现优于仅在SAYCam或同等规模通用数据上训练的模型。BabyVLM提供了一个发展对齐的评估工具,并展示了精心策划的小规模数据可使紧凑模型实现有效泛化,为数据高效视觉语言学习开辟新路径。

原文摘要 · Abstract (English)

Human infants rapidly develop visual reasoning skills from minimal input, suggesting that developmentally inspired pretraining could significantly enhance the efficiency of vision-language models (VLMs). Although recent efforts have leveraged infant-inspired datasets like SAYCam, existing evaluation benchmarks remain misaligned--they are either too simplistic, narrowly scoped, or tailored for large-scale pretrained models. Additionally, training exclusively on infant data overlooks the broader, diverse input from which infants naturally learn. To address these limitations, we propose BabyVLM, a novel framework comprising comprehensive in-domain evaluation benchmarks and a synthetic training dataset created via child-directed transformations of existing datasets. We demonstrate that VLMs trained with our synthetic dataset achieve superior performance on BabyVLM tasks compared to models trained solely on SAYCam or general-purpose data of the SAYCam size. BabyVLM thus provides a robust, developmentally aligned evaluation tool and illustrates how compact models trained on carefully curated data can generalize effectively, opening pathways toward data-efficient vision-language learning paradigms.

视觉语言模型数据效率婴儿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。