arXiv:2601.03400cs.CVcs.AI2026-01

构建多语言视觉谜题基准,测试模型深层推理能力

Eye-Q: A Multilingual Benchmark for Visual Word Puzzle Solving and Image-to-Phrase Reasoning

  • 设计隐含线索的视觉谜题,需抽象联想而非表面识别
  • 跨语言测试中最佳模型准确率仅60.27%,暴露推理短板
  • 适合评估视觉-语言模型在非字面语义映射上的表现

视觉-语言模型在标准基准上表现良好,但常依赖表层识别而非深层推理。本文提出视觉词谜题作为更具挑战性的评测方式,要求模型发现隐含视觉线索、生成并修正假设,将感知证据映射到非字面概念,难以通过字面锚定、OCR或简单检索解决。我们引入Eye-Q,一个包含1343个谜题的多语言基准,涵盖英语、波斯语、阿拉伯语及跨语言题目。每个谜题提供概念密集的场景与简短描述,要求模型推断特定目标词或短语。题目设计刻意无结构、线索隐晦,包含干扰项和上下文关系,需选择性注意、抽象与关联推理。采用开放式、人类对齐的评估协议,在轻量辅助下探测假设形成与修正过程。结果揭示显著性能差距,尤其在抽象与跨语言谜题上,表明当前模型在构建和搜索合适概念表示以实现灵活图像到短语推理方面存在局限,最高准确率为60.27%。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved strong performance on standard vision-language benchmarks, yet often rely on surface-level recognition rather than deeper reasoning. We propose visual word puzzles as a challenging alternative, as they require discovering implicit visual cues, generating and revising hypotheses, and mapping perceptual evidence to non-literal concepts in ways that are difficult to solve via literal grounding, OCR-heavy shortcuts, or simple retrieval-style matching. We introduce Eye-Q, a multilingual benchmark designed to assess this form of complex visual understanding. Eye-Q contains 1,343 puzzles in which a model observes a conceptually dense scene with a brief description and must infer a specific target word or phrase. The puzzles are intentionally unstructured and cue-implicit, with distractors and contextual relationships that demand selective attention, abstraction, and associative inference. The benchmark spans English, Persian, Arabic, and cross-lingual puzzles. We evaluate state-of-the-art VLMs using an open-ended, human-aligned protocol that probes hypothesis formation and revision under lightweight assistance. Results reveal substantial performance gaps, especially on abstract and cross-lingual puzzles, highlighting limitations in current models' ability to construct and search over appropriate conceptual representations for flexible image-to-phrase inference; maximum accuracy reaches only 60.27%.

视觉推理多语言词谜题图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。