arXiv:2505.21850cs.CVcs.AI2025-05ACL被引 5

提出多阶段抽象视觉推理新基准,揭示大模型在复杂规则识别中的短板

Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task

  • 基于RAVEN设计多阶段推理任务,模拟真实思维过程
  • 发现主流大模型在复杂规则检测上表现显著下降
  • 新指标MSEval可评估中间步骤正确性,更贴近人类推理

当前多模态大语言模型在一般视觉推理上表现优异,但在需要高层次思维的抽象视觉推理(AVR)方面研究不足。现有评测集中于单步推理,忽视了推理过程的多阶段特性。以往研究显示大模型在此类任务中表现不佳,但未说明具体失败原因。为此,我们提出MultiStAR——一个基于RAVEN的多阶段抽象视觉推理基准,用于评估不同复杂度下的推理能力。同时,传统准确率仅关注最终结果,忽略中间步骤。因此,我们引入MSEval新指标,综合考量中间步骤与最终输出的正确性。我们在17个代表性闭源与开源多模态大模型上进行实验,结果表明:尽管模型在基础感知任务中表现良好,但在复杂规则识别阶段仍面临显著挑战。

原文摘要 · Abstract (English)

Current Multimodal Large Language Models (MLLMs) excel in general visual reasoning but remain underexplored in Abstract Visual Reasoning (AVR), which demands higher-order reasoning to identify abstract rules beyond simple perception. Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process. Past studies found MLLMs struggle with these benchmarks, but it doesn't explain how they fail. To address this gap, we introduce MultiStAR, a Multi-Stage AVR benchmark, based on RAVEN, designed to assess reasoning across varying levels of complexity. Additionally, existing metrics like accuracy only focus on the final outcomes while do not account for the correctness of intermediate steps. Therefore, we propose a novel metric, MSEval, which considers the correctness of intermediate steps in addition to the final outcomes. We conduct comprehensive experiments on MultiStAR using 17 representative close-source and open-source MLLMs. The results reveal that while existing MLLMs perform adequately on basic perception tasks, they continue to face challenges in more complex rule detection stages.

抽象推理多模态模型评测基准大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。