arXiv:2511.03146cs.CL2025-11被引 3

构建视觉认知评测基准,揭示多模态模型在空间几何推理上的短板。

MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity

  • 按空间、几何、知识三类组织11项视觉推理任务
  • 闭源模型领先(如Gemini-2.5-Pro达42.66分),空间几何推理普遍低于30%
  • 发现方向误判、跨视角识别脆弱等共性错误,强调视觉提取关键作用

随着推理模型的快速发展,多模态在人类认知中的核心作用日益凸显,推动对以视觉为中心的认知行为进行深入评估。然而现有多模态基准或过度强调文本推理,或未能系统捕捉视觉认知行为,导致多模态大模型(MLLMs)的认知能力评估不足。为此,我们提出MME-CC(Multi-Modal Evaluation benchmark of Cognitive Capacity),一个以视觉为基础的评测基准,将11个代表性推理任务划分为空间、几何和知识三类,并对MLLMs在这些维度上的认知能力进行细粒度分析。基于MME-CC,我们在16个代表性MLLM上开展全面实验。结果表明,闭源模型整体领先(如Gemini-2.5-Pro达42.66分,优于GLM-4.5V的30.45分),但空间与几何推理仍普遍薄弱(均≤30%)。我们进一步识别出常见错误模式,包括方向判断失误、跨视角身份保持脆弱以及对反事实指令响应差,并观察到思维链通常遵循‘提取→推理→验证’三阶段流程,且严重依赖视觉提取。本工作旨在推动将多模态大模型的认知能力作为评估与设计的核心。

原文摘要 · Abstract (English)

As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behaviors. Yet, existing multimodal benchmarks either overemphasize textual reasoning or fall short of systematically capturing vision-centric cognitive behaviors, leaving the cognitive capacity of MLLMs insufficiently assessed. To address this limitation, we introduce MME-CC (Multi-Modal Evaluation benchmark of Cognitive Capacity), a vision-grounded benchmark that organizes 11 representative reasoning tasks into three fundamental categories of visual information: spatial, geometric, and knowledge-based reasoning, and provides fine-grained analyses of MLLMs' cognitive capacity across these dimensions. Based on MME-CC, we conduct extensive experiments over 16 representative MLLMs. Our study reveals that closed-source models currently lead overall (e.g., 42.66 for Gemini-2.5-Pro vs. 30.45 for GLM-4.5V), while spatial and geometric reasoning remain broadly weak (less than or equal to 30%). We further identify common error patterns, including orientation mistakes, fragile cross-view identity persistence, and poor adherence to counterfactual instructions, and observe that Chain-of-Thought typically follows a three-stage process (extract -> reason -> verify) with heavy reliance on visual extraction. We hope this work catalyzes a shift toward treating the cognitive capacity of MLLMs as central to both evaluation and model design.

多模态认知评测视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。