发现多模态大模型缺乏人类与生俱来的基础认知能力。
Core Knowledge Deficits in Multi-Modal Language Models
- 构建12个核心认知概念的评测基准CoreCognition,评估模型表现。
- 230个模型在低级认知任务上表现差,且随规模扩大无改善。
- 提出概念劫持法,揭示模型依赖捷径学习而非真正理解。
尽管多模态大语言模型(MLLMs)在高层感知和推理任务中表现出色,但在真实场景中的鲁棒性仍有限,常在对人类而言自然轻松的任务上表现不佳。我们假设这些缺陷源于缺乏核心知识——人类从幼儿期就具备的基本认知能力。为探究MLLMs中的核心知识表征,我们提出了CoreCognition,一个涵盖12个基于发展认知科学的核心知识概念的大规模基准。我们评估了230个模型,使用11种不同提示,共生成2,530个数据点进行分析。实验揭示四个关键发现,共同表明MLLMs存在核心知识缺失:它们在低层次能力上的表现持续低于高层次任务,且未表现出可扩展性。最后,我们提出概念劫持(Concept Hacking),一种新型受控评估方法,揭示随着模型规模增大,MLLMs并未向真正的核心知识理解演进,而是依赖捷径学习。
原文摘要 · Abstract (English)
While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge--rudimentary cognitive abilities innate to humans from early childhood. To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We evaluate 230 models with 11 different prompts, leading to a total of 2,530 data points for analysis. Our experiments uncover four key findings, collectively demonstrating core knowledge deficits in MLLMs: they consistently underperform and show reduced, or even absent, scalability on low-level abilities relative to high-level ones. Finally, we propose Concept Hacking, a novel controlled evaluation method that reveals MLLMs fail to progress toward genuine core knowledge understanding, but instead rely on shortcut learning as they scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。