arXiv:2604.01764cs.CV2026-04中稿 · ICLR

测试大模型破解文字谜题的能力,发现其认知推理严重不足。

Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning

  • 设计1164个文字谜题,检验视觉与语言知识融合推理能力。
  • 顶尖模型准确率低于10%,且增大模型或提示无法提升性能。
  • 适合研究多模态认知推理、具身智能的学者参考。

大型视觉语言模型(LVLMs)在显性视觉识别上表现优异,能准确描述图像中直接可见的内容。然而,当视觉输入仅作为线索而非答案时,模型表现出显著的认知缺陷。本文指出,当前模型难以完成需多步抽象推理的任务,如破解文字谜题。解决文字谜题需要模型提取视觉与文本属性,调用语言先验知识(如成语),并进行抽象映射,形成超越像素空间的语义。为此,我们提出RebusBench基准,包含1,164个谜题,用于评估该类神经符号推理能力。对Qwen、InternVL、LLaVA等先进模型的评测显示,其精确匹配率低于10%,语义准确率低于20%,且模型规模扩大或使用上下文学习(ICL)均未带来明显改进。这表明模型虽具备视觉与语言能力,但缺乏连接两者的认知推理机制。项目页面见https://amirkasaei.com/rebusbench/。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the visual input serves only as a clue rather than the answer. We identify that current models struggle with the complex, multi-step reasoning required to solve problems where information is not explicitly depicted. Successfully solving a rebus puzzle requires a distinct cognitive workflow: the model must extract visual and textual attributes, retrieve linguistic prior knowledge (such as idioms), and perform abstract mapping to synthesize these elements into a meaning that exists outside the pixel space. To evaluate this neurosymbolic capability, we introduce RebusBench, a benchmark of 1,164 puzzles designed to test this specific integration of perception and knowledge. Our evaluation of state-of-the-art models (including Qwen, InternVL, and LLaVA) shows a severe deficiency: performance saturates below 10% Exact Match and 20% semantic accuracy, with no significant improvement observed from model scaling or In-Context Learning (ICL). These findings suggest that while models possess the necessary visual and linguistic components, they lack the cognitive reasoning glue to connect them. Project page available at https://amirkasaei.com/rebusbench/.

视觉推理文字谜题认知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。