模型能看懂图像细节,却难理解讽刺意味,揭示视觉与理解间的鸿沟。
Are MLMs Trapped in the Visual Room?
- 用视觉房间类比,测试模型是否真理解而非仅识别图像。
- 8个顶尖模型对讽刺判断平均错17.1%,虽能看懂细节却不懂意图。
- 适合关注多模态理解深度、评估模型真实推理能力的研究者。
多模态大模型(MLMs)能否真正‘理解’图像?受塞尔的中文屋思想实验启发,我们提出‘视觉房间’论点:系统可按规则处理视觉输入的所有细节,却不具备真正的理解力。这挑战了‘感知即理解’的普遍假设。为此,我们构建了一个双层评估框架,涵盖感知与认知两个层面。感知部分评估模型对图像表面细节的捕捉能力,认知部分则考察其对讽刺极性的推断能力。为支持该框架,我们创建了一个高质量多模态讽刺数据集,包含924张静态图片和100段动态视频,所有讽刺标签均由原作者标注并经独立评审验证。我们评估了8个最先进的MLMs。结果表明:(1) 模型在视觉感知上表现优异;(2) 即便感知准确,其讽刺理解的平均错误率仍高达~17.1%,揭示‘看见’与‘理解’之间存在显著差距;(3) 此差距源于上下文整合、情感推理与语用推断能力的不足。本研究为视觉房间论提供了实证依据,并提出了新的多模态模型评估范式。
原文摘要 · Abstract (English)
Can multi-modal large models (MLMs) that can ``see'' an image be said to ``understand'' it? Drawing inspiration from Searle's Chinese Room, we propose the \textbf{Visual Room} argument: a system may process and describe every detail of visual inputs by following algorithmic rules, without genuinely comprehending the underlying intention. This dilemma challenges the prevailing assumption that perceptual mastery implies genuine understanding. In implementation, we introduce a two-tier evaluation framework spanning perception and cognition. The perception component evaluates whether MLMs can accurately capture the surface-level details of visual contents, where the cognitive component examines their ability to infer sarcasm polarity. To support this framework, We further introduce a high-quality multi-modal sarcasm dataset comprising both 924 static images and 100 dynamic videos. All sarcasm labels are annotated by the original authors and verified by independent reviewers to ensure clarity and consistency. We evaluate eight state-of-the-art (SoTA) MLMs. Our results highlight three key findings: (1) MLMs demonstrate high accuracy in visual perception; (2) even with correct perception, MLMs exhibit an average error rate of ~17.1\% in sarcasm understanding, revealing a significant gap between seeing and understanding; (3) this gap stems from weaknesses in context integration, emotional reasoning, and pragmatic inference. This work provides empirical grounding for the proposed Visual Room argument and offers a new evaluation paradigm for MLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。