arXiv:2510.02780cs.CV2025-10

通过解谜题揭示视觉语言模型的认知短板

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

  • 构建221个谜题数据集,按六类认知能力标注
  • 模型在图像组合上表现好,但对隐喻和文化符号理解差
  • 提示策略显著影响推理方式,解释力关乎性能

视觉语言模型(VLMs)在多模态任务中表现优异,但在复杂横向思维挑战(如字谜)中的认知过程仍不透明。尽管已有研究显示这些模型在字谜解答中表现不佳,其背后的推理机制与失败模式仍缺乏系统探索。本研究通过全面的可解释性分析,超越性能指标,深入理解VLM如何应对这类复杂思维挑战。我们构建了一个系统标注的221个字谜数据集,涵盖六类认知类别,并设计评估框架,将推理质量与答案正确性分离。研究考察三种提示策略,旨在激发不同类型的解释过程,揭示关键洞察:推理质量在不同谜题类别间差异显著,模型在视觉构图上具有系统性优势,但在缺席解读和文化象征理解上存在根本局限。同时发现,提示策略显著影响认知路径与求解效果,表明可解释性应作为模型性能的核心组成部分,而非事后补充。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel at many multimodal tasks, yet their cognitive processes remain opaque on complex lateral thinking challenges like rebus puzzles. While recent work has demonstrated these models struggle significantly with rebus puzzle solving, the underlying reasoning processes and failure patterns remain largely unexplored. We address this gap through a comprehensive explainability analysis that moves beyond performance metrics to understand how VLMs approach these complex lateral thinking challenges. Our study contributes a systematically annotated dataset of 221 rebus puzzles across six cognitive categories, paired with an evaluation framework that separates reasoning quality from answer correctness. We investigate three prompting strategies designed to elicit different types of explanatory processes and reveal critical insights into VLM cognitive processes. Our findings demonstrate that reasoning quality varies dramatically across puzzle categories, with models showing systematic strengths in visual composition while exhibiting fundamental limitations in absence interpretation and cultural symbolism. We also discover that prompting strategy substantially influences both cognitive approach and problem-solving effectiveness, establishing explainability as an integral component of model performance rather than a post-hoc consideration.

视觉语言模型可解释性认知分析字谜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。