构建视觉问答新基准,测试模型对抽象视觉的深层理解能力
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
- 设计4类风格化图像数据集,融合符号、文化与语言知识
- 仅Gemini-2.5-flash和GPT-4o表现良好,3.7%图像无模型识别
- 揭示模型在视觉隐喻、表情包等任务中的强弱差异
我们提出VisualQuest,一个新型数据集,用于严格评估多模态大模型(MLLMs)在需要整合符号、文化和语言知识的抽象视觉推理任务上的表现。不同于聚焦真实图像分类或直接描述的现有基准,VisualQuest包含3,551张非照片、风格化的图像,涵盖四大类别:公众人物、流行文化、语言表达与文学作品。每张图像配有针对性问题,以探测复杂推理能力。我们对十种顶尖MLLMs进行评测,发现仅有Gemini-2.5-flash和GPT-4o达到较强整体性能,而3.7%的图像未被任何模型识别,凸显多模态理解仍存重大挑战。细粒度分析表明,Gemini在识别风格化公众人物方面表现优异,而GPT-4o在视觉双关、表情符号组合等语言推理任务中领先。VisualQuest为推进抽象视觉推理研究提供全面且具挑战性的资源,并指明未来模型优化的关键方向。数据集已开源:https://github.com/xkt88/VISUALQUEST。
原文摘要 · Abstract (English)
We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration of symbolic, cultural, and linguistic knowledge. Unlike existing benchmarks that focus on direct image captioning or classification of realistic images, VisualQuest comprises 3,551 non-photographic, stylized images spanning four categories: Public Figures, Popular Culture, Linguistic Expressions, and Literary Works. Each image is paired with targeted questions to probe complex reasoning. We benchmark ten state-of-the-art MLLMs and find that only Gemini-2.5-flash and GPT-4o achieve strong overall performance, while 3.7 percent of the images remain unrecognized by any model, underscoring persistent challenges in multimodal understanding. Fine-grained analysis shows that Gemini excels at recognizing stylized public figures, whereas GPT-4o leads in linguistic reasoning tasks such as visual puns and emoji combinations. VisualQuest provides a comprehensive and challenging resource for advancing research in abstract visual reasoning and highlights key areas for future model improvement. The dataset is available at https://github.com/xkt88/VISUALQUEST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。