arXiv:2601.06521cs.CVcs.CL2026-01被引 28

测试大模型视觉推理能力,发现其远不如三岁小孩。

BabyVision: Visual Reasoning Beyond Language

  • 构建独立于语言的视觉能力评测基准BabyVision。
  • 顶级模型平均仅49.7分,不及六岁儿童。
  • 适合关注视觉本质理解与模型鲁棒性的研究者。

人类在掌握语言前已具备基本视觉能力,但当前多模态大模型仍严重依赖语言先验来弥补脆弱的视觉理解。我们发现:顶尖多模态大模型在人类三岁儿童也能轻松完成的基础视觉任务上表现持续失败。为系统探究这一差距,我们提出BabyVision,一个旨在评估多模态大模型核心视觉能力且不依赖语言知识的基准。该基准涵盖388个样本,分为22个子类,覆盖四大关键类别。实证结果与人工评估显示,领先模型性能显著低于人类基线:Gemini3-Pro-Preview得分为49.7,落后于六岁儿童表现,远低于成人平均分94.1。这表明,尽管在知识密集型测评中表现优异,当前多模态大模型仍缺乏基础视觉原语。在BabyVision上取得进展,是迈向人类级视觉感知与推理的重要一步。我们还提出了BabyVision-Gen及自动化评估工具链,探索生成式模型解决视觉推理的路径。代码与数据已开源至https://github.com/UniPat-AI/BabyVision。

原文摘要 · Abstract (English)

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.

视觉推理多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。