arXiv:2603.18472cs.AIcs.CV2026-03被引 6

发现多模态大模型在符号理解上存在感知与推理倒置现象。

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

  • 构建跨领域的符号理解基准,分三层认知层级评测
  • 模型识别基础符号能力差,复杂推理却表现较好
  • 适合关注视觉符号理解瓶颈的研究者和开发者

多模态大语言模型在自然图像上表现优异,但对离散视觉符号的理解能力尚不明确。本文构建了一个覆盖语言、文化、数学、物理和化学的多领域基准,分为感知识别、组合推理、关联批判三个认知层级。实验显示,主流多模态大模型普遍存在认知错配:在基础符号识别任务中表现不佳,而在更复杂的推理任务中反而相对稳健。这种识别-推理倒置现象表明,当前系统常依赖语言先验、模板检索或程序化推理,而非可靠的视觉根基。该模式在手写字符、公式图、电路图、化学结构等稀疏低冗余符号中尤为明显。结果表明,符号理解仍是多模态智能的重大瓶颈,亟需优先考虑离散语义空间中具身感知的训练与评估方案。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathematics, physics and chemistry, organized into three cognitive levels: perception and recognition, combination and reasoning, and association and critical thinking. Across leading MLLMs, we observe a consistent cognitive mismatch. Models frequently underperform on elementary symbol recognition while appearing relatively competent on more complex reasoning tasks. This recognition-reasoning inversion indicates that current systems often compensate with linguistic priors, template retrieval or procedural reasoning instead of robust visual grounding. The pattern is especially clear for sparse, low-redundancy symbols such as handwritten characters, formula graphs, circuit diagrams and chemical structures. These results show that symbolic understanding remains a major bottleneck for multimodal intelligence and motivate training and evaluation schemes that prioritize grounded perception in discrete semantic spaces.

符号理解多模态认知评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。