arXiv:2601.08292cs.CV2026-01被引 1

测试大模型像6岁孩子一样看图,发现差距巨大。

KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?

  • 用儿童视觉能力拆解视觉任务,构建新基准KidVis
  • 顶级模型仅达人类儿童67.33分,平均差28分
  • 参数越大越不提升基础视觉能力,存在悖论

尽管多模态大语言模型在高级推理任务中表现优异,但其是否具备与6-7岁儿童相当的基础视觉感知能力仍不确定。为此,我们提出KidVis,一个基于人类视觉发展理论的新基准。该基准将视觉智能分解为六种原子能力:注意力、追踪、区分、记忆、空间、闭合性,涵盖10类低语义依赖的视觉任务。评估20个前沿多模态大模型发现,人类儿童平均得分高达95.32,而最先进的GPT-5仅得67.33。关键的是,我们观察到‘缩放定律悖论’:单纯增加模型参数无法线性提升这些基础视觉能力。研究证实,当前多模态大模型虽具推理能力,却缺乏通用视觉智能所必需的生理级感知原语。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have demonstrated impressive proficiency in high-level reasoning tasks, such as complex diagrammatic interpretation, it remains an open question whether they possess the fundamental visual primitives comparable to human intuition. To investigate this, we introduce KidVis, a novel benchmark grounded in the theory of human visual development. KidVis deconstructs visual intelligence into six atomic capabilities - Concentration, Tracking, Discrimination, Memory, Spatial, and Closure - already possessed by 6-7 year old children, comprising 10 categories of low-semantic-dependent visual tasks. Evaluating 20 state-of-the-art MLLMs against a human physiological baseline reveals a stark performance disparity. Results indicate that while human children achieve a near-perfect average score of 95.32, the state-of-the-art GPT-5 attains only 67.33. Crucially, we observe a "Scaling Law Paradox": simply increasing model parameters fails to yield linear improvements in these foundational visual capabilities. This study confirms that current MLLMs, despite their reasoning prowess, lack the essential physiological perceptual primitives required for generalized visual intelligence.

多模态视觉理解认知评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。