arXiv:2502.00698cs.AIcs.CV2025-02被引 24

提出首个评估多模态模型抽象与推理能力的基准MM-IQ

MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

  • 构建涵盖8种推理范式的4776道视觉推理题训练集
  • 顶尖模型准确率仅33.17%,远低于人类水平
  • 适合关注AI认知能力评估的研究者

智力测试是评估人类认知能力的基础方法,通过剥离语言背景、语言水平或领域知识,专注于抽象与推理的核心能力。然而,当前人工智能研究缺乏系统性的基准来量化多模态系统在这些关键认知能力上的表现。为此,我们提出MM-IQ,一个综合性评估框架,包含大规模训练集(4,776个视觉推理问题)和2,710个精心设计的测试题,覆盖8种不同推理范式。对现有开源及专有多模态模型的系统评估显示,即使最先进的架构,准确率也仅略高于随机猜测(33.17%对比25%基线),显著差距凸显当前模型在模拟人类基本推理能力上的不足,亟需范式变革以弥合这一认知鸿沟。此外,受大推理模型兴起启发,我们还发布了一个基于可验证奖励函数的强化学习训练的多模态推理模型,以更小规模实现接近顶尖性能。

原文摘要 · Abstract (English)

IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or domain-specific knowledge to isolate core competencies in abstraction and reasoning. Yet, artificial intelligence research currently lacks systematic benchmarks to quantify these critical cognitive capabilities in multimodal systems. To address this crucial gap, we propose MM-IQ, a comprehensive evaluation framework, which comprises a large-scale training set with 4,776 visual reasoning problems and 2,710 meticulously curated test items spanning 8 distinct reasoning paradigms. Through systematic evaluation of existing open-source and proprietary multimodal models, our benchmark reveals striking limitations: even state-of-the-art architectures achieve only marginally superior performance to random chance (33.17% vs. 25% baseline accuracy). This substantial performance chasm highlights the inadequacy of current multimodal models in approximating fundamental human reasoning capacities, underscoring the need for paradigm-shifting advancements to bridge this cognitive divide. Moreover, inspired by the recent surge of large reasoning models, we also release a multimodal reasoning model as the baseline that is trained via reinforcement learning with verifiable reward functions, reaching competitive performance to the state-of-the-art with a notably smaller model size.

多模态推理评估认知能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。