arXiv:2504.10342cs.CL2025-04被引 46

新基准VisualPuzzles专测视觉推理,避开专业知识干扰。

VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

  • 用中文公考逻辑题改编,设计五类纯推理问题
  • 顶尖多模态模型在该基准上仍远落后于人类表现
  • 适合评估真实推理能力,不依赖知识记忆

现有多模态评测常将推理与特定领域知识混淆,难以独立评估非专家场景下的通用推理能力。为此,我们提出VisualPuzzles,一个聚焦视觉推理且刻意减少专业背景依赖的基准。该基准包含五类问题:算法、类比、演绎、归纳和空间推理,主要来源为中文公务员考试中的逻辑题人工翻译。实验表明,VisualPuzzles对领域知识要求显著低于MMMU等基准,更强调复杂推理,能更好评估真实的多模态推理能力。评测显示,当前最先进多模态大模型在该任务上持续落后于人类,且在知识密集型评测中表现好,并不意味着在轻知识、重推理任务上也成功。此外,提升推理计算量(如“思考”模式)对不同模型和任务效果不一,模型规模与性能间无明确关联。模型在VisualPuzzles上的推理和作答模式也与知识导向型基准存在差异。该基准为评估超越事实记忆与领域知识的推理能力提供了更清晰视角。

原文摘要 · Abstract (English)

Current multimodal benchmarks often conflate reasoning with domain-specific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings. To address this, we introduce VisualPuzzles, a benchmark that targets visual reasoning while deliberately minimizing reliance on specialized knowledge. VisualPuzzles consists of diverse questions spanning five categories: algorithmic, analogical, deductive, inductive, and spatial reasoning. One major source of our questions is manually translated logical reasoning questions from the Chinese Civil Service Examination. Experiments show that VisualPuzzles requires significantly less intensive domain-specific knowledge and more complex reasoning compared to benchmarks like MMMU, enabling us to better evaluate genuine multimodal reasoning. Evaluations show that state-of-the-art multimodal large language models consistently lag behind human performance on VisualPuzzles, and that strong performance on knowledge-intensive benchmarks does not necessarily translate to success on reasoning-focused, knowledge-light tasks. Additionally, reasoning enhancements such as scaling up inference compute (with "thinking" modes) yield inconsistent gains across models and task types, and we observe no clear correlation between model size and performance. We also found that models exhibit different reasoning and answering patterns on VisualPuzzles compared to benchmarks with heavier emphasis on knowledge. VisualPuzzles offers a clearer lens through which to evaluate reasoning capabilities beyond factual recall and domain knowledge.

多模态推理评测基准知识解耦视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。