arXiv:2505.12000cs.CV2025-05被引 4

用真人智商题测试视觉语言模型的真推理能力,发现模型答对了但想错了。

IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests

  • 设计纯视觉智商题集,避免文字线索干扰模型
  • 三款顶尖模型平均准确率最高仅61.5%,3D推理几乎全错
  • 重点评估解题思路而非答案,揭示模型“伪推理”问题

尽管大型视觉语言模型在多模态任务中表现优异,其在人类智商测试中的真实推理能力仍待深入研究。为推动该领域进展,我们提出IQBench,一个针对标准化视觉智商测试的新基准。该基准以视觉为核心,最大限度减少对文本内容的依赖,促使模型主要从图像信息中推导答案,而非依赖预训练的文本知识。为此,我们人工收集并标注了500道视觉智商题,防止训练中无意的数据泄露。不同于以往仅关注最终答案准确率的研究,我们通过评估模型解释、解题模式、最终预测准确率及人工评价来综合衡量其推理能力。实验表明,不同任务间性能差异显著:o4-mini、gemini-2.5-flash、claude-3.7-sonnet平均准确率分别为0.615、0.578、0.548;所有模型在3D空间与字谜推理任务上均表现不佳,暴露出当前视觉语言模型普遍存在的泛化推理局限。在推理得分方面,三者平均分分别为0.696、0.586、0.516。结果表明,模型的最终答案与推理过程之间存在明显不一致,强调在评估时需同时关注推理准确性与输出一致性。

原文摘要 · Abstract (English)

Although large Vision-Language Models (VLMs) have demonstrated remarkable performance in a wide range of multimodal tasks, their true reasoning capabilities on human IQ tests remain underexplored. To advance research on the fluid intelligence of VLMs, we introduce **IQBench**, a new benchmark designed to evaluate VLMs on standardized visual IQ tests. We focus on evaluating the reasoning capabilities of VLMs, which we argue are more important than the accuracy of the final prediction. **Our benchmark is visually centric, minimizing the dependence on unnecessary textual content**, thus encouraging models to derive answers primarily from image-based information rather than learned textual knowledge. To this end, we manually collected and annotated 500 visual IQ questions to **prevent unintentional data leakage during training**. Unlike prior work that focuses primarily on the accuracy of the final answer, we evaluate the reasoning ability of the models by assessing their explanations and the patterns used to solve each problem, along with the accuracy of the final prediction and human evaluation. Our experiments show that there are substantial performance disparities between tasks, with models such as `o4-mini`, `gemini-2.5-flash`, and `claude-3.7-sonnet` achieving the highest average accuracies of 0.615, 0.578, and 0.548, respectively. However, all models struggle with 3D spatial and anagram reasoning tasks, highlighting significant limitations in current VLMs' general reasoning abilities. In terms of reasoning scores, `o4-mini`, `gemini-2.5-flash`, and `claude-3.7-sonnet` achieved top averages of 0.696, 0.586, and 0.516, respectively. These results highlight inconsistencies between the reasoning processes of the models and their final answers, emphasizing the importance of evaluating the accuracy of the reasoning in addition to the final predictions.

视觉语言模型智商测试推理能力评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。