构建首个评估视觉横向思维的多选题基准,发现模型表现远逊于人类。
COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
- 将视觉横向思维转化为多选题问答任务,基于成语和常见短语生成谜题。
- 创建包含1000+个谜题的合成数据集,每题4个选项,测试模型抽象联想能力。
- 揭示当前视觉语言模型在自动生成抽象描述上存在明显缺陷,适合研究认知推理者。
尽管视觉问答(VQA)基准推动了推理技术发展,但主要聚焦于纵向思维。有效解决问题还需横向思维,这一能力在人工智能中尚未得到充分研究,也未用于检验视觉感知系统。为此,我们提出将视觉横向思维定义为多选题问答任务,并设计三步分类驱动的方法论来构建任务实例。进而,开发了COLUMBUS——一个合成基准,通过该任务流程,基于公开的复合词和常见短语集合生成文本与图标谜题。COLUMBUS包含超过1000个谜题,每个谜题有四个答案选项。尽管当前最先进的视觉语言模型(VLMs)表现尚可,但评估显示人类与模型之间存在显著差距。模型虽能受益于人工标注的描述,但在自主生成恰当抽象表示方面仍显不足。
原文摘要 · Abstract (English)
While visual question-answering (VQA) benchmarks have catalyzed the development of reasoning techniques, they have focused on vertical thinking. Effective problem-solving also necessitates lateral thinking, which remains understudied in AI and has not been used to test visual perception systems. To bridge this gap, we formulate visual lateral thinking as a multiple-choice question-answering task and describe a three-step taxonomy-driven methodology for instantiating task examples. Then, we develop COLUMBUS, a synthetic benchmark that applies the task pipeline to create QA sets with text and icon rebus puzzles based on publicly available collections of compounds and common phrases. COLUMBUS comprises over 1,000 puzzles, each with four answer candidates. While the SotA vision-language models (VLMs) achieve decent performance, our evaluation demonstrates a substantial gap between humans and models. VLMs benefit from human-curated descriptions but struggle to self-generate such representations at the right level of abstraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。