arXiv:2605.19130cs.LGcs.AI2026-05

评测视觉语言模型在自然手摄视频中的跨模态学习能力

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

论文配图:EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
图 1 · 摘自论文原文
  • 用不同语义对齐程度的数据训练模型,测试其在真实手摄视频上的表现
  • 现有模型依赖强对齐数据,在弱对齐信号上表现差,远低于人类婴儿水平
  • 提出EgoBabyVLM挑战赛,推动从自然手摄数据中学习语言的模型发展

儿童在有限的视听输入下能以惊人鲁棒性习得语言意义,远超当前最优的大规模多模态模型。近期研究指出,基于网络精选数据训练的视觉语言模型(VLMs)难以泛化到可穿戴设备、具身代理和婴儿头戴相机产生的稀疏、弱对齐手摄视频流——且目前缺乏固定评估流程来衡量该场景下的进展。我们使用包含不同语义对齐程度的视觉-语言数据集训练VLMs,涵盖自然手摄的婴幼儿与成人视频,并通过涵盖多模态语言接地及单模态视觉与语言任务的综合评测套件进行评估。核心是Machine-DevBench,一个基于语料库的词汇与语法能力评测,自动从模型训练词汇的对数频率分箱中生成,以消除先前发展基准的训练/测试不匹配与统计效力不足问题。结果表明,当前VLM范式严重依赖精选数据的紧密语义对齐,无法有效利用主导自然手摄输入的弱对齐信号——而这正是人类擅长的领域。为推动进展,我们提出EgoBabyVLM挑战赛,激励开发能够从人类婴儿所经历的自然手摄数据中实现语言扎根学习的模型。

原文摘要 · Abstract (English)

Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs) trained on curated web data fail to generalize to the sparse, weakly-aligned egocentric streams produced by wearable devices, embodied agents, and infant head-cams -- and no fixed evaluation pipeline exists for measuring progress on this regime. We train VLMs on datasets with varying degrees of semantic alignment between visual and linguistic inputs, including naturalistic infant and adult egocentric videos, and evaluate them with a comprehensive suite spanning multimodal language grounding and unimodal vision and language tasks. At the core of this suite is Machine-DevBench, a corpus-grounded benchmark of lexical and grammatical competence, automatically generated from the model's training vocabulary across logarithmic frequency bins to eliminate the train/eval mismatch and low statistical power of prior developmental benchmarks. Our results show that current VLM paradigms hinge on the tight semantic alignment of curated data and fail to exploit the weakly-aligned signal that dominates naturalistic egocentric input -- the very regime in which humans thrive. To motivate progress, we introduce the EgoBabyVLM Challenge to drive the development of models capable of grounded language learning from the kind of naturalistic data that human infants experience.

多模态学习手摄视频语言接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。