arXiv:2505.11141cs.CVcs.AI2025-05被引 2

构建细粒度人类对齐基准,评估多模态模型推理能力与人类差距

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans

  • 设计包含9794个上下文推理题的多模态基准,覆盖四类推理任务
  • 发现当前多模态大模型在推理准确率上显著低于人类表现
  • 提供人类错误选项和正确率数据,助力模型改进方向定位

实现通用人工智能(AGI)的目标是模仿并超越人类。OpenAI的o1、o3以及DeepSeek的R1等模型已展现出类人推理能力,性能优异,并逐步融入多模态大语言模型(MLLMs)。然而,这些模型在处理推理任务时是否具备与人类相当的能力仍不明确。本文提出Human-Aligned Bench,一个用于多模态推理与人类表现精细对齐的基准。我们收集了9,794个仅依赖上下文推理的多模态问题,包括中英文双语多模态题和纯文本题,涵盖视觉推理、定义判断、类比推理和逻辑判断四类。每个问题均附带人类正确率及易错选项。在该基准上的大量实验揭示了当前多模态大模型在多模态推理能力上与人类存在显著差距。研究结果为下一代模型的发展提供了重要启示。

原文摘要 · Abstract (English)

The goal of achieving Artificial General Intelligence (AGI) is to imitate humans and surpass them. Models such as OpenAI's o1, o3, and DeepSeek's R1 have demonstrated that large language models (LLMs) with human-like reasoning capabilities exhibit exceptional performance and are being gradually integrated into multimodal large language models (MLLMs). However, whether these models possess capabilities comparable to humans in handling reasoning tasks remains unclear at present. In this paper, we propose Human-Aligned Bench, a benchmark for fine-grained alignment of multimodal reasoning with human performance. Specifically, we collected 9,794 multimodal questions that solely rely on contextual reasoning, including bilingual (Chinese and English) multimodal questions and pure text-based questions, encompassing four question types: visual reasoning, definition judgment, analogical reasoning, and logical judgment. More importantly, each question is accompanied by human success rates and options that humans are prone to choosing incorrectly. Extensive experiments on the Human-Aligned Bench reveal notable differences between the performance of current MLLMs in multimodal reasoning and human performance. The findings on our benchmark provide insights into the development of the next-generation models.

多模态推理人类对齐模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。