arXiv:2506.11080cs.CL2025-06ACL被引 1

评测多模态模型是否真比人聪明,发现它在深层推理上仍落后。

MANBench: Is Your Multimodal Model Smarter than Human?

  • 构建中英双语基准,覆盖9类任务1314个问题。
  • 模型在图文理解上超人,但跨模态推理仍弱于人类。
  • 适合研究多模态智能与人类认知差距的学者使用。

多模态大语言模型(MLLMs)的快速发展引发了关于其能否超越人类在多模态任务中表现的讨论。为此,我们提出MANBench(多模态能力基准),一个包含1314个问题的中英文双语基准,覆盖九项任务,涵盖基于知识与非知识领域。MANBench强调直观推理、跨模态无缝整合及现实复杂性,提供严谨评估框架。通过涉及多样参与者的大量人类实验,我们将人类表现与前沿MLLMs进行对比。结果表明,尽管MLLMs在知识型和文本-图像理解任务中表现优异,但在深层跨模态推理任务如转换理解、图像一致性及多图理解方面仍存在明显短板。此外,人类与MLLMs在高复杂度任务如谜题和空间想象中均面临挑战。MANBench揭示了当前模型的优劣势,显示先进模型在多数领域仍未达到人类水平表现。我们希望MANBench能推动弥合模型与人类多模态能力之间的差距。代码与数据集已公开于https://github.com/micdz/MANBench。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has ignited discussions regarding their potential to surpass human performance in multimodal tasks. In response, we introduce MANBench (Multimodal Ability Norms Benchmark), a bilingual benchmark (English and Chinese) comprising 1,314 questions across nine tasks, spanning knowledge-based and non-knowledge-based domains. MANBench emphasizes intuitive reasoning, seamless cross-modal integration, and real-world complexity, providing a rigorous evaluation framework. Through extensive human experiments involving diverse participants, we compared human performance against state-of-the-art MLLMs. The results indicate that while MLLMs excel in tasks like Knowledge and Text-Image Understanding, they struggle with deeper cross-modal reasoning tasks such as Transmorphic Understanding, Image Consistency, and Multi-image Understanding. Moreover, both humans and MLLMs face challenges in highly complex tasks like Puzzles and Spatial Imagination. MANBench highlights the strengths and limitations of MLLMs, revealing that even advanced models fall short of achieving human-level performance across many domains. We hope MANBench will inspire efforts to bridge the gap between MLLMs and human multimodal capabilities. The code and dataset are available at https://github.com/micdz/MANBench.

多模态模型评测人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。