测试视觉语言模型解谜能力,发现它们难理解隐喻和抽象联想。
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
- 构建英文回文谜题数据集,涵盖图像、空间布局与语言双关。
- 模型能解简单图像替换,但对抽象推理和文字游戏基本失效。
- 适合研究多模态理解瓶颈或人类认知启发的AI设计者。
回文谜题通过图像、空间关系和符号替代编码语言,对当前视觉语言模型(VLMs)构成独特挑战。不同于传统图像描述或问答任务,解谜需跨模态抽象、符号推理及对文化、语音和语言双关的理解。本文构建了一个人工生成并标注的多样化英文回文谜题基准,涵盖从简单象形替换到依赖空间线索(如“头”在“脚”之上)的复杂形式。我们分析了多种VLM的表现,结果表明尽管模型在解析简单视觉线索方面展现一定能力,但在需要抽象推理、发散思维和视觉隐喻理解的任务上表现显著不足。
原文摘要 · Abstract (English)
Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs). Unlike traditional image captioning or question answering tasks, rebus solving requires multi-modal abstraction, symbolic reasoning, and a grasp of cultural, phonetic and linguistic puns. In this paper, we investigate the capacity of contemporary VLMs to interpret and solve rebus puzzles by constructing a hand-generated and annotated benchmark of diverse English-language rebus puzzles, ranging from simple pictographic substitutions to spatially-dependent cues ("head" over "heels"). We analyze how different VLMs perform, and our findings reveal that while VLMs exhibit some surprising capabilities in decoding simple visual clues, they struggle significantly with tasks requiring abstract reasoning, lateral thinking, and understanding visual metaphors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。